> ## Content Index
> Fetch the complete content index at: https://roundproxies.com/blog/llms.txt
> Use this file to discover other available public pages before exploring further.

# Web Scraping with Gologin: Step-by-Step Guide (2026)
- URL: https://roundproxies.com/blog/web-scraping-gologin/
- Published: 2026-01-27T22:37:15.000Z
- Updated: 2026-09-24T15:36:25.000Z
- Description: Learn web scraping with GoLogin. Step-by-step guide with Selenium, Playwright, and Puppeteer code examples. Bypass bot detection in 2026.
- Author: Marius Bernard
- Tags: Web Scraping, #dated-69788aed26f439f88a95b550, Golang (Go)

Web scraping with Gologin earns its price at one specific moment: your scraper parses pages fine, but the site's anti-bot layer flags the browser within a few requests.

New IPs don't fix that. The canvas hash, WebGL renderer, and navigator values travel with you to every IP you rotate through.

Gologin gives each scraper session its own browser profile with a separate fingerprint, cookie jar, and proxy. You drive it with the Playwright, Selenium, or Puppeteer code you already have.

Older guides still copy the deprecated `gl.create()` call and a pricing table from before January 2026\. This one uses the SDK's current methods.

Every snippet runs against books.toscrape.com, a sandbox built for scraping practice.

You'll also get a pre-flight check that catches proxy and timezone mismatches before they burn a profile, plus fixes for the actual error messages Gologin throws.

## What is web scraping with Gologin?

Web scraping with Gologin means driving Gologin's Orbita browser profiles with Playwright, Selenium, or Puppeteer so each scraper session carries its own browser fingerprint, cookies, and proxy. The Python SDK starts a profile and returns a debugger address; your automation tool connects to it over CDP. Use it when stealth plugins stop passing fingerprint checks.

Gologin is an antidetect browser. Its engine, Orbita, is a modified Chromium that reports the fingerprint stored in the profile instead of your machine's real values.

The profile itself lives on Gologin's servers. When you call `start()`, the SDK downloads it, launches Orbita on your machine, and hands back a `host:port` for the Chrome DevTools Protocol.

After that, Gologin steps aside. Selectors, waits, and pagination are ordinary Playwright or Selenium code.

## What Gologin fixes and what stays your problem

Antidetect browsers get marketed as block-proof. They handle one detection layer well and leave the others to you, so it pays to know which is which before you spend money.

| Detection layer     | What the site checks                          | Gologin handles it? | What you still own                |
| ------------------- | --------------------------------------------- | ------------------- | --------------------------------- |
| Browser fingerprint | Canvas, WebGL, audio, fonts, navigator values | Yes, per profile    | Keep OS and user agent consistent |
| Automation traces   | navigator.webdriver, CDP side effects         | Partly              | Verify it in Step 5               |
| IP reputation       | ASN type, IP history, blocklists              | No                  | Decent proxies                    |
| Geo consistency     | Timezone and language vs. IP location         | Partly              | Re-check after proxy changes      |
| Behavior            | Click paths, timing, scroll patterns          | No                  | Delays, realistic navigation      |
| Rate limits         | Requests per IP or session                    | No                  | Throttling, spreading load        |

If a Gologin scraper still gets blocked, check the rows marked "No" first. A perfect fingerprint behind a flagged datacenter IP loses to a basic IP reputation check.

## Prerequisites: the free plan won't run this code

This trips up more readers than anything else. Gologin's Forever Free plan gives you three profiles, but [its own pricing docs](https://gologin.com/docs/general/account-and-billing/pricing) list API access as unavailable on it.

The SDK authenticates every call with an API token, so you need the 7-day trial or a paid plan to follow along.

You'll also need:

- A current Python 3 release and a virtual environment
- A Gologin API token (dashboard: Settings, then API, then "New Token")
- Proxy credentials with HTTP or SOCKS5 auth
- On a Linux server with no display: Xvfb, which Gologin's server docs recommend for running Orbita

## Step 1: Install the SDK and store your token

Install the official SDK (the PyPI package is `gologin`, from the [pygologin repo](https://github.com/gologinapp/pygologin)) plus the automation libraries. Keep the token in an environment variable so it never lands in Git.

```bash
python -m venv .venv && source .venv/bin/activate
pip install gologin playwright selenium webdriver-manager

# Keep secrets out of your code
export GL_API_TOKEN="paste-your-token-here"

```

You can skip `playwright install`. Playwright attaches to Orbita over CDP, so it never launches its own bundled Chromium. Drop `selenium` and `webdriver-manager` if you're Playwright-only.

## Step 2: Create a profile you'll reuse

Create profiles once and store their IDs. A profile is closer to a long-lived user account than a browser tab, and its value grows as it collects cookies and history.

```python
import os
from gologin import GoLogin

gl = GoLogin({"token": os.environ["GL_API_TOKEN"]})

# Pulls a real-device fingerprint set for the OS you choose
profile = gl.createProfileRandomFingerprint({"os": "win", "name": "books-scraper-01"})
profile_id = profile["id"]

print(profile_id)  # save this in your config or .env as GL_PROFILE_ID

```

Stick with `"win"` unless your target's audience skews toward Macs. Linux desktops are a small slice of consumer traffic, so a Linux fingerprint stands out on retail sites.

Resist the urge to create a fresh profile per run. Each one counts toward your plan's profile limit, and a brand-new profile shows up with zero cookie history.

If you need a fixed language or screen size, `createProfileWithCustomParams` accepts a `navigator` block. The SDK README lists every field it takes.

## Step 3: Attach your own proxy

Gologin's fingerprint does nothing for IP reputation, so pair each profile with its own proxy. `changeProfileProxy` writes the proxy into the stored profile.

```python
gl.changeProfileProxy(profile_id, {
    "mode": "http",            # or "socks5"
    "host": "proxy.example.com",
    "port": 8000,
    "username": os.environ["PROXY_USER"],
    "password": os.environ["PROXY_PASS"],
})

```

Use sticky sessions. A rotating gateway that hands one profile a new IP on every request looks like one person teleporting between cities, which is exactly the inconsistency fingerprinting systems look for.

On protected retail and ticketing sites, residential or ISP IPs hold up far better than datacenter ranges.

Roundproxies sells both with sticky sessions. Any provider with HTTP or SOCKS5 auth plugs into the dict above the same way.

Gologin also has `addGologinProxyToProfile(profile_id, "us")`, which draws on the 2 GB of residential traffic bundled with paid plans. That covers a trial run, not a production crawl loading full pages with images.

## Step 4: Start the profile and connect Playwright

`gl.start()` downloads the profile, launches Orbita, and returns a debugger address like `127.0.0.1:35421`. Playwright attaches to it with `connect_over_cdp`.

```python
from playwright.sync_api import sync_playwright

gl = GoLogin({"token": os.environ["GL_API_TOKEN"], "profile_id": profile_id})
debugger_address = gl.start()

try:
    with sync_playwright() as p:
        browser = p.chromium.connect_over_cdp(f"http://{debugger_address}")
        context = browser.contexts[0]   # the profile's own context, cookies included
        page = context.pages[0] if context.pages else context.new_page()
        page.goto("https://books.toscrape.com/", wait_until="domcontentloaded")
        print(page.title())
finally:
    gl.stop()  # saves cookies and state back to the profile

```

Use `browser.contexts[0]`. Calling `browser.new_context()` creates a blank context with none of the profile's cookies or storage, which throws away the thing you're paying Gologin for.

Keep `gl.stop()` in the `finally` block. Skipping it leaves orphaned Orbita processes, which on Windows show up later as the `EBUSY` error covered below.

## Step 5: Run a pre-flight check

This is the step most Gologin tutorials skip. If the proxy exits in Frankfurt but the profile reports `America/New_York`, the fingerprint is internally inconsistent, and detection vendors score that.

The mismatch creeps in when you swap proxies on an existing profile or use a gateway that hops countries. One request catches it.

```python
import json

def preflight(page):
    """Compare what the proxy says about us with what the browser says."""
    page.goto("https://ipinfo.io/json")          # goes out through Orbita's proxy
    ip = json.loads(page.locator("body").inner_text())
    env = page.evaluate("""() => ({
        tz: Intl.DateTimeFormat().resolvedOptions().timeZone,
        lang: navigator.language,
        webdriver: navigator.webdriver,
    })""")
    problems = []
    if env["tz"] != ip.get("timezone"):
        problems.append(f"timezone {env['tz']} vs proxy {ip.get('timezone')}")
    if env["webdriver"]:
        problems.append("navigator.webdriver is true")
    return ip.get("ip"), problems

```

Check the IP by navigating the page, as above. Calls made through `page.request` are sent by Playwright itself rather than the browser, so they don't reliably inherit the proxy configured inside Orbita.

The script covers the cheap checks. Once per new profile, open it by hand and run a full fingerprint audit with [CreepJS](https://roundproxies.com/blog/creepjs/).

Then test for [WebRTC leaks](https://roundproxies.comblog/webrtc-leaks/), which can expose your real IP even with the proxy working.

## Step 6: Scrape and paginate

With the profile connected, the extraction code is plain Playwright. This function pulls title, price, and stock status from one listing page.

```python
def scrape_books(page):
    rows = []
    for card in page.locator("article.product_pod").all():
        rows.append({
            "title": card.locator("h3 a").get_attribute("title"),  # full title lives here
            "price": card.locator(".price_color").inner_text(),
            "in_stock": "In stock" in card.locator(".availability").inner_text(),
            "url": page.url,
        })
    return rows

```

The visible link text on books.toscrape.com is truncated with an ellipsis. The `title` attribute holds the full name, which is why the selector reads the attribute.

Pagination follows the "next" link and yields rows page by page, so the caller can save them immediately.

```python
import random
import time
from urllib.parse import urljoin

def crawl(page, url, max_pages=5):
    for _ in range(max_pages):
        page.goto(url, wait_until="domcontentloaded")
        yield from scrape_books(page)
        nxt = page.locator("li.next a")
        if nxt.count() == 0:
            break
        url = urljoin(page.url, nxt.get_attribute("href"))  # hrefs are relative
        time.sleep(random.uniform(2, 5))  # pace like a reader

```

`urljoin` matters here. The first "next" link is `catalogue/page-2.html` and later ones are `page-3.html`, so string concatenation breaks on page two.

## Step 7: Save rows as you go

Append each row to a JSON Lines file the moment you have it. If Orbita crashes on page 40, you keep the first 39.

```python
def save(row, path="books.jsonl"):
    with open(path, "a", encoding="utf-8") as f:
        f.write(json.dumps(row, ensure_ascii=False) + "\n")

```

The full run below wires the pieces together. It refuses to scrape if the pre-flight check finds a problem.

```python
def main():
    gl = GoLogin({"token": os.environ["GL_API_TOKEN"],
                  "profile_id": os.environ["GL_PROFILE_ID"]})
    debugger_address = gl.start()
    try:
        with sync_playwright() as p:
            browser = p.chromium.connect_over_cdp(f"http://{debugger_address}")
            page = browser.contexts[0].new_page()
            ip, problems = preflight(page)
            if problems:
                raise SystemExit(f"Profile unsafe on {ip}: {problems}")
            for row in crawl(page, "https://books.toscrape.com/", max_pages=3):
                save(row)
    finally:
        gl.stop()

if __name__ == "__main__":
    main()

```

Three pages on this site gives you 60 rows. Point `crawl` at your real target once the selectors match.

## Gologin with Selenium

Selenium attaches to the same debugger address through `debuggerAddress`. The one trap is the driver version: it has to match Orbita's Chromium, not the Chrome installed on your machine.

```python
from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager

debugger_address = gl.start()
version = gl.get_chromium_version()  # Orbita's version, not your system Chrome

options = webdriver.ChromeOptions()
options.add_experimental_option("debuggerAddress", debugger_address)
driver = webdriver.Chrome(
    service=Service(ChromeDriverManager(driver_version=version).install()),
    options=options,
)
driver.get("https://books.toscrape.com/")

```

When you're done, call `driver.quit()`, wait a second or two, then `gl.stop()`. Stopping Gologin first can leave the driver holding a dead session.

For Gologin work, I'd pick Playwright: auto-waiting, and no driver version to keep in sync with Orbita.

If you're still deciding for a new project, the [Playwright vs Selenium comparison](https://roundproxies.com/blog/playwright-vs-selenium/) covers waits and speed in detail.

## Gologin with Puppeteer in Node.js

The Node SDK runs Orbita locally too, so you don't need Gologin's paid cloud browser for Puppeteer. `start()` returns a WebSocket URL that `puppeteer.connect` accepts.

```javascript
// scrape.mjs  (ES module, so top-level await works)
import GoLogin from 'gologin';
import puppeteer from 'puppeteer-core';

const GL = new GoLogin({
  token: process.env.GL_API_TOKEN,
  profile_id: process.env.GL_PROFILE_ID,
});

const { wsUrl } = await GL.start();
const browser = await puppeteer.connect({ browserWSEndpoint: wsUrl.toString(), defaultViewport: null });
try {
  const page = await browser.newPage();
  await page.goto('https://books.toscrape.com/', { waitUntil: 'domcontentloaded' });
  const titles = await page.$$eval('article.product_pod h3 a', (els) => els.map((e) => e.title));
  console.log(`${titles.length} books`, titles.slice(0, 3));
} finally {
  await browser.disconnect(); // detach, don't kill Orbita
  await GL.stop();            // Orbita shuts down and syncs the profile
}

```

Install with `npm i gologin puppeteer-core`. Use `puppeteer-core` rather than `puppeteer`, since you're connecting to Orbita and don't need a second bundled Chromium.

Setting `defaultViewport: null` keeps the profile's own screen size. Puppeteer's default 800x600 viewport would clash with the resolution in the fingerprint.

## Scaling web scraping with Gologin: a profile pool

The expensive part of a Gologin scraper is profile startup: an API round trip, a profile download, and an Orbita launch.

Starting a profile per URL repeats that cost every time and eats into your plan's API rate limit (300 requests per minute on Professional).

The pattern that scales is a small pool of long-lived profiles, each with a sticky proxy, each handling a batch of URLs per session.

### Batch URLs per profile

Each worker starts one profile, processes its batch, and stops. First, a helper that asks the OS for an unused port.

```python
import socket

def free_port():
    with socket.socket() as s:
        s.bind(("127.0.0.1", 0))  # port 0 = "give me any free port"
        return s.getsockname()[1]

```

Passing that port to the SDK keeps two Orbita instances from fighting over one debugging port when workers run side by side.

```python
def worker(profile_id, urls):
    gl = GoLogin({"token": os.environ["GL_API_TOKEN"],
                  "profile_id": profile_id, "port": free_port()})
    addr = gl.start()
    try:
        with sync_playwright() as p:
            page = p.chromium.connect_over_cdp(f"http://{addr}").contexts[0].new_page()
            for url in urls:
                page.goto(url, wait_until="domcontentloaded")
                for row in scrape_books(page):
                    save(row, path=f"books-{profile_id}.jsonl")  # one file per worker
                time.sleep(random.uniform(2, 5))
    finally:
        gl.stop()

```

One output file per profile avoids interleaved writes from parallel processes. Merge them when the run finishes.

### Run profiles in parallel

Processes beat threads here. Each worker owns its own Orbita, and a crash in one doesn't take the others down.

```python
from concurrent.futures import ProcessPoolExecutor

PROFILE_IDS = ["profile_a", "profile_b", "profile_c"]

def run(urls):
    n = len(PROFILE_IDS)
    chunks = [urls[i::n] for i in range(n)]  # spread URLs evenly across profiles
    with ProcessPoolExecutor(max_workers=n) as pool:
        futures = [pool.submit(worker, pid, chunk) for pid, chunk in zip(PROFILE_IDS, chunks)]
    for f in futures:
        f.result()  # re-raise worker errors instead of swallowing them

if __name__ == "__main__":
    run([f"https://books.toscrape.com/catalogue/page-{n}.html" for n in range(1, 31)])

```

Budget memory before you raise the worker count. Gologin's own server guidance puts each running profile at roughly 300 to 500 MB of RAM, and heavy target pages push that higher.

### Retire profiles that start failing

Track blocks per profile: 403s, CAPTCHA pages, or a page title that doesn't match what you expected. Once a profile trips repeatedly on a target, stop using it there.

You can give a slot a fresh identity without deleting it:

```python
gl = GoLogin({"token": os.environ["GL_API_TOKEN"]})
gl.refreshProfilesFingerprint([burned_id])     # new fingerprint, same profile slot
gl.changeProfileProxy(burned_id, fresh_proxy)  # and a new exit IP to match

```

The old cookies still belong to the previous identity. If the site tied the block to a session cookie, delete the profile with `gl.delete(burned_id)` and create a clean one.

## Headless mode

Gologin passes Chromium flags through `extra_params`, which is how its own docs enable headless mode.

```python
gl = GoLogin({
    "token": os.environ["GL_API_TOKEN"],
    "profile_id": profile_id,
    "extra_params": ["--headless"],
})

```

On strict targets I'd run headed Orbita under Xvfb instead. Headless Chromium has rendering and API differences that some detection scripts probe.

A virtual display removes that variable for a little extra RAM.

## Gologin pricing for scrapers in 2026

Gologin changed its plans in January 2026\. These are the published prices for new accounts at the time of writing; check the [official pricing page](https://gologin.com/docs/general/account-and-billing/pricing) before you buy.

| Plan         | Profiles         | REST API limit | Cloud launches | Monthly   | Annual (per month) |
| ------------ | ---------------- | -------------- | -------------- | --------- | ------------------ |
| Forever Free | 3                | No API access  | None           | $0        | $0                 |
| Professional | 10, 50, or 100   | 300 RPM        | 1 concurrent   | From $9   | From $4.50         |
| Business     | 300 or 500       | 500 RPM        | 2 concurrent   | From $119 | From $59.50        |
| Enterprise   | 1,000            | 800 RPM        | 3 concurrent   | $299      | $149.50            |
| Custom       | 2,000 to 100,000 | 1,200 RPM      | 4 concurrent   | From $449 | From $224.50       |

Every paid plan includes 2 GB of Gologin residential proxy traffic. Accounts registered before January 2026 stay on legacy pricing unless you ask support to move them.

For the local-Orbita approach in this tutorial, the profile count and API limit are the numbers that matter. Cloud launch limits only apply if you run profiles on Gologin's servers.

## When Gologin is the wrong tool

Gologin is a monthly bill on top of your proxy bill. Before you pay it, rule out the cheaper options.

If the target doesn't fingerprint browsers at all, an HTTP client with good headers beats any browser on speed and cost.

If you're blocked on behavior or rate limits, a better fingerprint won't change the outcome.

For moderate protection, open-source stealth browsers often do the job for free:

| Tool       | Cost              | How it handles fingerprints                      | Pick it when                                      |
| ---------- | ----------------- | ------------------------------------------------ | ------------------------------------------------- |
| Gologin    | Paid plan         | Stored real-device profiles, persistent sessions | You need many persistent identities with cookies  |
| Camoufox   | Free, open source | Firefox build that spoofs at the engine level    | You want stealth without a subscription           |
| Patchright | Free, open source | Patched Playwright that removes automation leaks | Your Playwright code only needs the leaks plugged |
| nodriver   | Free, open source | Drives Chrome over CDP with no WebDriver layer   | You're in Python and want a light setup           |

I'd try [Camoufox](https://roundproxies.com/blog/camoufox/) first on a new target. Move to Gologin once you need persistent logged-in sessions or dozens of stable identities, since the free tools don't manage those for you.

Gologin isn't the only antidetect browser with an API, either. The [MoreLogin automation guide](https://roundproxies.com/blog/morelogin-web-automation/) follows the same connect-over-CDP pattern if you want to compare.

## Troubleshooting Gologin errors

These are the messages you'll hit most, with causes from Gologin's [common errors reference](https://gologin.com/docs/general/troubleshooting/common-errors).

### "Profile already deleted"

**Why:** Misleading wording. It usually means you've hit the profile launch limit on a trial or free plan; the profile still exists. **Fix:** Upgrade or renew the plan. Gologin keeps profiles on its servers for up to 180 days after a subscription lapses.

### "ECONNREFUSED 127.0.0.1"

**Why:** Orbita failed to start, so nothing is listening on the debugger port. Common causes are a corrupted Orbita download or a proxy that stops the browser from initializing. **Fix:** Launch the profile once with the proxy removed. If it starts, the proxy is the problem; if not, force a fresh Orbita download from the app settings.

### "SyntaxError: Unexpected token 'e'"

**Why:** Cloudflare in front of Gologin's API rate-limited you (error 1015) and returned HTML where the client expected JSON. A JSON decode error from a Python SDK call points to the same cause. **Fix:** Slow down profile starts and add backoff.

```python
def start_with_backoff(gl, attempts=4):
    for i in range(attempts):
        try:
            return gl.start()
        except Exception as exc:  # the SDK raises generic errors on API hiccups
            wait = 5 * 2 ** i     # 5s, 10s, 20s, 40s
            print(f"start() failed ({exc}); retrying in {wait}s")
            time.sleep(wait)
    raise RuntimeError("Profile would not start after retries")

```

Swap `gl.start()` for `start_with_backoff(gl)` in the worker. Staggering worker startup by a few seconds helps too.

### "EBUSY: resource busy or locked"

**Why:** Windows only. A previous Orbita process didn't shut down and still holds a lock on the profile's cookie file. **Fix:** End every `orbita` process in Task Manager, then delete that profile's folder under `AppData\Local\GoLogin\browser\orbita\profiles\`. Server-side profile data isn't affected.

### "session not created: This version of ChromeDriver only supports Chrome version ..."

**Why:** Selenium downloaded a driver for your system Chrome instead of Orbita's Chromium. **Fix:** Pass `gl.get_chromium_version()` into `ChromeDriverManager(driver_version=...)`, as in the Selenium section.

### "The proxy is functional but not working on your device or network"

**Why:** Gologin verified the proxy from its servers, but your network or firewall blocks the connection from your machine. **Fix:** Allow the proxy port through your firewall, or test from another network to confirm.

## FAQ

### Is Gologin free for web scraping?

Not for automated scraping. The Forever Free plan has three profiles but no API access, which the SDK needs. Use the 7-day trial to test your scripts.

### Can Gologin bypass Cloudflare?

Gologin covers the fingerprint part of Cloudflare's scoring. Cloudflare also scores IP reputation and behavior, so a profile on a flagged datacenter IP can still get challenged.

### How many Gologin profiles can I run at once?

Locally, RAM sets the limit: Gologin estimates 300 to 500 MB per running profile. Cloud launches have hard caps per plan, from one concurrent session on Professional to four on Custom.

### Does Gologin work with Playwright?

Yes. Start the profile with the SDK, then call `chromium.connect_over_cdp()` with the returned address. Use `browser.contexts[0]` so you keep the profile's cookies.

### Is scraping with Gologin legal?

The tool is legal; the risk depends on what you collect and how. Public data at a polite rate is lower risk than logged-in areas or personal data. Respect robots.txt, and get legal advice for commercial projects.

## Wrapping up

Web scraping with Gologin comes down to four habits: reuse profiles, pin one sticky proxy to each, verify the profile before trusting it, and always call `stop()`.

Start with one profile and the `main()` script from Step 7\. Add the profile pool once that profile scrapes your real target cleanly for a full session; scaling a leaky setup gets you blocked faster.