Selenium vs. Playwright for scraping in 2026

Both tools drive a browser. The choice barely matters until you put a proxy behind it and try to run fifty sessions at once.

Most comparisons ranking for this query answer a testing question: which one has better assertions, better reporting, better CI integration. Useful if you write end-to-end tests. Useless if you're scraping.

This one answers the scraping question. How each tool handles proxy rotation, proxy authentication, request blocking and detection, with the code and the failure modes.

Selenium vs. Playwright: the core difference

Selenium talks to a browser through a separate driver binary over HTTP, while Playwright holds one WebSocket open to a browser it ships itself. For scraping, that means Playwright gives you a proxy per browser context and native request interception, where Selenium gives you one browser per IP and CDP workarounds.

Selenium vs. Playwright at a glance

Selenium is the better choice when your language, your browser or your existing Grid rules out Playwright.

Ruby, Kotlin, real Safari on real hardware, an Appium device farm: Selenium reaches places Playwright doesn't. A working Grid beats a 30% speedup.

Playwright is the better choice for almost everything else in scraping. Contexts make per-IP isolation cheap and route handlers make bandwidth control trivial.

Auto-waiting deletes a category of bug you'd otherwise write yourself.

Dimension Selenium Playwright
Proxy rotation ⭐⭐ One browser process per IP ⭐⭐⭐⭐⭐ One context per IP, same process
Proxy auth ⭐⭐ Needs an extension or IP whitelisting ⭐⭐⭐⭐⭐ Username and password fields in the API
Request blocking ⭐⭐⭐ CDP, Chromium only, version-pinned ⭐⭐⭐⭐⭐ page.route() on all three engines
Waiting ⭐⭐ Explicit waits, written by you ⭐⭐⭐⭐⭐ Actionability checks on every action
Browser reach ⭐⭐⭐⭐⭐ Anything with a WebDriver ⭐⭐⭐ Chromium, Firefox, WebKit
Language bindings ⭐⭐⭐⭐⭐ Java, Python, C#, Ruby, JS, more ⭐⭐⭐ JS/TS, Python, Java, .NET
Stealth ecosystem ⭐⭐⭐ Active, but mostly by abandoning WebDriver ⭐⭐⭐⭐ Patchright and rebrowser, both current
Distributed runs ⭐⭐⭐⭐⭐ Grid is mature and boring ⭐⭐⭐ Roll your own orchestration

Playwright waits for elements to be clickable; Selenium doesn't

Selenium's find_element returns as soon as a node is in the DOM. Whether it's visible, whether something is sitting on top of it, whether it finished animating: not Selenium's problem.

So you write the waiting logic yourself, and every scraper ends up with its own half-correct version of it.

from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By

# Without this, the click lands on whatever was under the cursor.
wait = WebDriverWait(driver, 15)
wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, "#load-more"))).click()

That works. It also has to be repeated at every interaction, and the version someone writes at 1am usually isn't element_to_be_clickable, it's time.sleep(3).

Playwright runs the same checks before every action, without being asked:

# Waits for #load-more to be attached, visible, stable and enabled.
page.click("#load-more")

On a server-rendered page this barely matters.

On a React listing where the node exists a few hundred milliseconds before it's clickable, it's the difference between a scraper that finishes and one that silently collects empty pages.

Selenium needs one browser per IP; Playwright needs one context

This is where the scraping answer diverges from the testing answer.

Chrome takes its proxy from a command-line flag at launch. A Selenium driver is one Chrome process, so one driver is one exit IP, for its whole life.

from selenium import webdriver

def driver_for(host_port):
    opts = webdriver.ChromeOptions()
    opts.add_argument(f"--proxy-server=http://{host_port}")
    opts.add_argument("--headless=new")
    return webdriver.Chrome(options=opts)

# Ten IPs means ten Chrome processes, each with its own memory footprint.
drivers = [driver_for(hp) for hp in ("1.2.3.4:8000", "5.6.7.8:8000")]

Ten IPs, ten browser launches, ten sets of renderer processes. Concurrency is capped by RAM long before it's capped by your proxy pool.

Playwright attaches the proxy to a BrowserContext instead, and contexts are cheap. One browser process, many exits:

from playwright.sync_api import sync_playwright

GATEWAYS = ["http://gw1.example.com:8000", "http://gw2.example.com:8000"]

with sync_playwright() as p:
    # The "per-context" placeholder is required on Chromium/Windows.
    browser = p.chromium.launch(proxy={"server": "per-context"})
    for gw in GATEWAYS:
        ctx = browser.new_context(proxy={"server": gw})
        page = ctx.new_page()
        page.goto("https://httpbin.org/ip")
        print(page.inner_text("body"))
        ctx.close()
    browser.close()

The placeholder at launch is the part people miss. Skip it and per-context proxies fall back to the browser-level setting on Chromium under Windows.

It's in Playwright's network guide and almost nowhere else.

There's a Selenium-shaped workaround: point Chrome at 127.0.0.1 and run a local forwarder that rewrites the upstream on each connection.

You keep one browser and rotate underneath it. You also now maintain a proxy server, so be sure the RAM you're saving covers that.

Chrome drops your proxy password

--proxy-server has nowhere to put credentials. Pass http://user:pass@gw:8000 and Chrome discards the user info, then throws a native auth dialog that WebDriver can't click. Your requests hang or come back 407.

The fix is a small MV3 extension that sets the proxy and answers the auth challenge. The whole thing is two files, and this is the half that matters:

// background.js
chrome.proxy.settings.set({
  value: { mode: "fixed_servers",
           rules: { singleProxy: { scheme: "http", host: "gw.example.com", port: 8000 } } },
  scope: "regular"
});

chrome.webRequest.onAuthRequired.addListener(
  () => ({ authCredentials: { username: "USER", password: "PASS" } }),
  { urls: ["<all_urls>"] },
  ["blocking"]
);

The manifest needs "permissions": ["proxy", "webRequest", "webRequestAuthProvider"] and "host_permissions": ["<all_urls>"]. That third permission is what still allows a blocking onAuthRequired listener under Manifest V3; without it the listener loads and does nothing.

Load it with opts.add_argument("--load-extension=/path/to/ext"), and use --headless=new, since old headless ignores extensions entirely. There's more on the failure cases in our Selenium proxy setup guide.

Playwright takes the credentials as fields:

ctx = browser.new_context(proxy={
    "server": "http://gw.example.com:8000",
    "username": "USER",
    "password": "PASS",
})

One gotcha survives here too: Playwright ignores credentials embedded in the server URL. They have to be separate keys, or you get the same 407 with no useful error.

What a per-context proxy does not isolate

This is the part that catches people who migrate to Playwright for the rotation and assume the job is done.

A context isolates storage. Cookies, localStorage, IndexedDB, service workers, cache: all separate, which is exactly what you want for multi-account work.

It does not isolate the machine. Every context in that browser reports the same canvas hash, the same WebGL renderer, the same font list and the same TLS fingerprint.

There's one process on one host, and the host answers honestly.

Five contexts on five residential IPs aren't five users. They're one machine appearing from five places at once, and a detector that stores fingerprints links them to each other in a single query.

Timezone is the loudest version of this. The context carries the locale and timezone you set, or the one from launch, so an IP in São Paulo paired with America/New_York is a free signal you handed over.

Set timezone_id and locale per context to match the exit. When a target fingerprints hard, give each identity its own browser process.

Clean residential exits from any provider, ours included, won't fix a shared canvas hash.

Playwright blocks images natively; Selenium needs CDP

The cheapest speedup in browser scraping is refusing to download things you'll never parse. Playwright puts that in the API:

def keep_only_data(route):
    if route.request.resource_type in ("document", "xhr", "fetch", "script"):
        route.continue_()
    else:
        route.abort()

page.route("**/*", keep_only_data)

Drop script from that tuple and you'll break most SPAs, so check what the target needs before getting aggressive. Our guide to intercepting network requests in Playwright goes further.

Selenium's equivalent goes through CDP:

driver.execute_cdp_cmd("Network.enable", {})
driver.execute_cdp_cmd("Network.setBlockedURLs",
                       {"urls": ["*.png", "*.jpg", "*.woff2", "*.css"]})

It works, with conditions. Chromium only, so no Firefox. URL patterns only, so no filtering by resource type.

And it's pinned to a DevTools version: Selenium 4.39 shipped CDP support for v143, v142 and v141, which tells you how that treadmill runs.

Selenium is migrating this surface to WebDriver BiDi, and recent releases are mostly BiDi work. Network interception through BiDi will close the gap. It isn't closed yet.

Selenium's best stealth tools aren't really Selenium

The honest summary is uncomfortable for Selenium: the strongest anti-detection options in its family work by leaving WebDriver behind.

undetected-chromedriver's author moved on to nodriver, which drives Chrome over CDP with no chromedriver at all.

SeleniumBase's CDP Mode is actively maintained, with patches landing through 2026 chasing Chrome releases, and it also bypasses the driver for the sensitive parts.

Both are good. Neither is quite Selenium anymore.

Meanwhile selenium-stealth still pulls roughly 294,000 downloads a month on a release from November 2020. Its default signature claims to be Chrome 83.

Anything comparing that against a live Chrome version flags it instantly.

Playwright's side patches the library rather than replacing it. Patchright avoids the Runtime.enable CDP call that exposes an attached DevTools client, strips --enable-automation and adds --disable-blink-features=AutomationControlled.

rebrowser-patches fixes the same leak independently, which is a decent signal the fix is the right one.

Both are narrow by their own documentation. Neither changes your canvas, your fonts or your TLS handshake.

Our breakdowns of Patchright in practice and SeleniumBase UC Mode cover what each one reaches.

Benchmark it on your own targets

Published Selenium vs. Playwright numbers range from "13% faster" to "2x faster," and both can be true, because the gap depends almost entirely on how much waiting your target forces.

On a server-rendered page where the data ships in the initial HTML, the transport difference is close to noise. On a hydration-heavy SPA, where the polling loop dominates, it's large.

Any number you're quoted was measured on pages that aren't yours.

Fifteen lines settles it for your case:

import time
from playwright.sync_api import sync_playwright

URLS = ["https://your-target.example/page/1",
        "https://your-target.example/page/2"]

with sync_playwright() as p:
    browser = p.chromium.launch()
    page = browser.new_page()
    t0 = time.perf_counter()
    for url in URLS:
        page.goto(url, wait_until="domcontentloaded")
        page.wait_for_selector(".product-card")  # your real extraction anchor
    print(f"{(time.perf_counter() - t0) / len(URLS):.2f}s/page")
    browser.close()

Run the Selenium equivalent with the same wait condition and the same URLs. Keep the anchor identical in both, or you're timing two different jobs and calling it a comparison.

When Selenium wins

Your language isn't covered. Ruby and Kotlin teams have no Playwright binding, and rewriting a working scraper in TypeScript to save a few hundred milliseconds per page is a bad trade.

You already run Grid. Node autoscaling, queueing and session routing are solved there. Rebuilding that around Playwright costs more engineering time than the speed buys back.

You need real browsers on real hardware. Playwright ships patched Chromium, Firefox and WebKit builds. Selenium drives the retail Chrome your users have, plus Safari, plus Appium's device farms.

The target is server-rendered and the scrape is small. A few thousand pages of plain HTML, no hydration, no interception: the architecture argument mostly evaporates, and your existing code is the asset.

When Playwright wins

You're rotating IPs hard. Per-context proxies plus millisecond context creation is the single biggest practical difference in this comparison.

Bandwidth is the bill. Route handlers cut image, font and media traffic on every engine, without CDP version-matching.

You're memory-bound. Contexts share one browser process. The same box runs far more concurrent sessions than one-Chrome-per-IP allows.

You're starting fresh. No drivers to version-match against Chrome's release train, and auto-waiting means the flaky-click bug never gets written.

Proxy errors you'll hit in both tools

407 Proxy Authentication Required in Selenium means Chrome ate your credentials. Use the extension above, or whitelist your server IP with the provider and drop the credentials entirely.

net::ERR_TUNNEL_CONNECTION_FAILED is the proxy refusing the CONNECT, usually wrong port, dead exit, or a gateway that expects HTTPS on a different port than HTTP. Test the endpoint with curl before blaming either library.

Error: net::ERR_NO_SUPPORTED_PROXIES in Playwright almost always means a SOCKS5 proxy with credentials. Chromium doesn't support authenticated SOCKS5, so switch that gateway to HTTP.

Per-context proxy silently ignored on Chromium under Windows means you skipped proxy={"server": "per-context"} at launch. Every context then exits from the browser-level proxy, and your rotation does nothing.

StaleElementReferenceException in Selenium is the page re-rendering between your lookup and your click. Re-find inside the retry, don't cache the element. Playwright locators re-resolve on each use, which is why the exception has no equivalent there.

The verdict

For scraping, I'd take Playwright, and the deciding factor isn't speed. It's that proxy rotation, request blocking and session isolation are API surface instead of workarounds.

Selenium gets there. The authenticated-proxy extension works, CDP blocking works, and BiDi is closing the interception gap release by release.

You're maintaining scaffolding the other tool hands you for free.

Stay on Selenium when your language, browser matrix or existing Grid makes the choice for you. Those are constraints rather than sentiment, and a working scraper beats a faster architecture you haven't built yet.

Whichever you pick, the browser is the second decision. Check whether the site exposes JSON first, because no browser at all beats both tools by an order of magnitude.

Our guide to scraping with headless browsers covers the cases where a browser earns its keep.

FAQ

Is Playwright faster than Selenium for scraping?

Usually, and the margin depends on the target. On JavaScript-heavy pages the WebDriver round trip per command dominates, so the gap is wide. On server-rendered HTML it narrows enough that it shouldn't drive a migration.

Can Selenium rotate proxies without restarting the browser?

Not natively. Chrome reads --proxy-server at launch, so a new IP means a new driver. The workaround is pointing Chrome at a local forwarder you control and changing the upstream there.

Does Playwright get blocked less than Selenium?

Out of the box, both get blocked. Playwright doesn't set navigator.webdriver the way a driver-attached Chrome does, but it leaks CDP signals instead. Different surfaces, similar outcome.

Is Selenium being replaced by Playwright?

Not in enterprise QA, where Grid and W3C compliance keep it entrenched, and Selenium 4.49 shipped in September 2026 with a Selenium 5 charter published. In scraping specifically, new projects mostly start on Playwright.

Can I use both in the same project?

Yes, and it's a reasonable migration path. Playwright can attach to a running Chrome over CDP, including one Selenium launched, so you can borrow route interception without rewriting your existing driver code.