Most Zillow scraping tutorials end the same way. The BeautifulSoup code returns a 403, and the author suggests you pay for their API.

This guide shows how to scrape Zillow with code you run and control: six Python methods, easiest first, no paid scraping services.

You'll read the JSON Zillow hides in every page, call its internal search endpoint, and get past both PerimeterX blocks and the 500-result cap.

How do you scrape Zillow?

To scrape Zillow, request a search or property page with a browser-grade HTTP client like curl_cffi, then parse the JSON in the page's __NEXT_DATA__ script tag instead of the HTML. For bulk results, call Zillow's internal search endpoint directly. Add residential proxies and slower pacing once blocks start.

Zillow is a Next.js app, so the server ships every listing on the page as one JSON blob before React renders anything.

That blob changes far less often than class names like StyledPropertyCardDataArea-c11n-8-84-3, which get regenerated on every design-system deploy. Parse the JSON and a redesign won't break you.

What data can you scrape from Zillow?

You can scrape anything Zillow shows a logged-out visitor. Search results carry the summary fields; property pages carry history, valuations, and the full facts table.

Data Where it lives Method
Price, address, beds, baths, sqft, coordinates, status Search listResults / mapResults 1, 2, 6
zpid (Zillow property ID) and detail URL Search results 1, 2, 6
Zestimate and Rent Zestimate Search hdpData.homeInfo (sometimes), property page 3
Price history and tax history Property page gdpClientCache 3
Year built, lot size, HOA, parking, heating Property page resoFacts 3
Listing agent name and phone Property page 3 (treat as personal data)

Zillow renames keys every so often. Print .keys() on your first run before you hard-code any path from this article.

How does Zillow block scrapers?

Zillow runs HUMAN Security's bot defense, the product formerly called PerimeterX. It scores every request and answers low scores with a 403 page asking you to "Press & Hold" a button.

Four signals do most of the scoring:

  • TLS fingerprint. Python requests and httpx open TLS connections differently from Chrome, and the server sees that before it reads a single header.
  • IP reputation. Datacenter ranges get flagged fast; residential and ISP addresses look like house hunters.
  • Browser behavior. A sensor script collects timing, mouse movement, and browser properties. Sending no sensor data at all is its own signal.
  • Request rate. Hundreds of hits per minute from one IP.

Separately, Zillow caps results per query: the map returns at most 500 homes, and the list stops after about 20 pages. Method 6 works around that.

Methods overview

Methods 1, 2, 3, and 6 are plain HTTP requests through curl_cffi. Methods 4 and 5 add a real browser for when Zillow stops trusting plain HTTP.

Method Difficulty Cost Reliability Best for
1. Hidden JSON on search pages Easy Free Medium Quick checks, first ~40 listings
2. Internal search endpoint Easy Free Medium Paginated and filtered searches
3. Property detail pages Easy Free Medium Price history, Zestimate, facts
4. In-page fetch with Playwright Medium Free High When curl_cffi gets blocked
5. Stealth browser + cookie handoff Hard Proxy bill High Sustained daily volume
6. Map tiling Medium Free High Every listing in a city

"Free" means no software cost. Past a few hundred requests a day, every method needs residential or ISP proxies.

Prerequisites

You need Python 3.10 or newer, which curl_cffi has required since v0.14, plus three packages and their browsers:

pip install curl_cffi playwright patchright
playwright install chromium   # for Method 4
patchright install chrome     # for Method 5

Methods 2, 3, and 6 reuse the session and helpers from Method 1.

Method 1: Parse the hidden JSON on Zillow search pages

Method 1 fetches a search page with curl_cffi and reads the listings straight out of the __NEXT_DATA__ script tag.

Best when: you want one page of results fast, or you're checking whether your IP still gets through.

curl_cffi copies Chrome's TLS fingerprint, which gets you past the first check that kills requests.

This block sets up the session and pulls out the JSON with a regex, so you skip the HTML parser.

import json
import random
import re
import time
from curl_cffi import Session

session = Session(impersonate="chrome")  # Chrome TLS + HTTP/2 fingerprint
session.headers.update({"accept-language": "en-US,en;q=0.9"})

NEXT_DATA = re.compile(r'<script[^>]*id="__NEXT_DATA__"[^>]*>(.*?)</script>', re.S)

def get_next_data(url: str) -> dict:
    resp = session.get(url, timeout=30)
    match = NEXT_DATA.search(resp.text)
    if resp.status_code != 200 or not match:
        # PerimeterX block pages don't include a __NEXT_DATA__ tag
        raise RuntimeError(f"Blocked or layout changed: HTTP {resp.status_code}")
    return json.loads(match.group(1))

If this raises on your first request, run pip install -U curl_cffi before blaming the code. Stale Chrome fingerprints get flagged.

Next, a small helper that finds a key anywhere in the JSON tree. Zillow nests listResults several levels deep, and the nesting shifts between deploys.

def find_key(obj, key):
    """Depth-first search for the first value stored under `key`."""
    if isinstance(obj, dict):
        if key in obj:
            return obj[key]
        children = obj.values()
    elif isinstance(obj, list):
        children = obj
    else:
        return None
    for child in children:
        found = find_key(child, key)
        if found is not None:
            return found
    return None

data = get_next_data("https://www.zillow.com/brooklyn-new-york-ny/")
listings = find_key(data, "listResults") or []
for home in listings[:5]:
    print(home.get("zpid"), home.get("price"), home.get("address"))

Page 1 holds about 40 listings. Appending 2_p/ to the URL gets page 2, but every HTML page costs far more bandwidth than the JSON call in Method 2.

Tradeoff: no JavaScript runs. Once PerimeterX wants sensor data from your IP, this method gets block pages no matter how clean your headers look.

Method 2: Call Zillow's internal search endpoint

Method 2 sends the same PUT request Zillow's own frontend sends to /async-create-search-page-state, and gets clean JSON back with pagination included.

Best when: you need every page of a search, or filtered searches like rentals and recently sold homes.

Don't hand-build the payload. Load a search URL once and copy its queryState; every filter you set in the browser rides along.

SEARCH_API = "https://www.zillow.com/async-create-search-page-state"

def put_search(body: dict) -> dict:
    resp = session.put(
        SEARCH_API,
        json=body,
        timeout=30,
        headers={"origin": "https://www.zillow.com", "referer": "https://www.zillow.com/"},
    )
    resp.raise_for_status()
    return resp.json()

def search_page(query_state: dict, page: int = 1) -> dict:
    qs = {**query_state, "pagination": {"currentPage": page}}
    return put_search({
        "searchQueryState": qs,
        "wants": {"cat1": ["listResults", "mapResults"], "cat2": ["total"]},
        "requestId": page + 1,
    })

A browser sends origin and referer on these calls, so leaving them off makes you easier to spot.

Now seed queryState from a real search page and walk the pages with random pauses.

seed = get_next_data("https://www.zillow.com/brooklyn-new-york-ny/rentals/")
query_state = seed["props"]["pageProps"]["searchPageState"]["queryState"]

first = search_page(query_state)
total_pages = first["cat1"]["searchList"]["totalPages"]
homes = first["cat1"]["searchResults"]["listResults"]

for page in range(2, min(total_pages, 20) + 1):
    time.sleep(random.uniform(3, 7))  # human-ish pacing between pages
    data = search_page(query_state, page)
    homes += data["cat1"]["searchResults"]["listResults"]

print(f"{len(homes)} rental listings across {total_pages} pages")

The seed URL sets the filter. Use /rentals/ for rentals, or apply the Sold filter in your browser and paste the resulting URL; the queryState carries it either way.

Tradeoff: the list stops at about 20 pages, roughly 800 homes. For a whole metro, move to Method 6.

Method 3: Scrape Zillow property pages for price history and Zestimates

Method 3 loads individual property pages and decodes the listing from gdpClientCache, a JSON string stored inside the page's JSON.

Best when: you need the full record for specific homes: price history, tax history, Zestimate, and the facts table.

The URL comes from any search result's detailUrl field, so this chains onto Method 2.

from urllib.parse import urljoin

def scrape_property(detail_url: str) -> dict:
    data = get_next_data(urljoin("https://www.zillow.com", detail_url))
    cache = find_key(data, "gdpClientCache")
    if cache is None:
        raise KeyError("No gdpClientCache: check for a block page or a new layout")
    entries = json.loads(cache)  # second decode: the cache is a JSON string
    return next(iter(entries.values()))["property"]

prop = scrape_property(homes[0]["detailUrl"])
print(prop.get("price"), prop.get("zestimate"), prop.get("rentZestimate"))
for event in prop.get("priceHistory", [])[:5]:
    print(event.get("date"), event.get("event"), event.get("price"))

Some property pages keep the data in a hdpApolloPreloadedData script tag instead. If gdpClientCache is missing and you didn't get a block page, look for that tag.

Detail pages are where personal data lives: agent names, phone numbers, and sometimes owner names on for-sale-by-owner listings. Store only what your project needs.

Tradeoff: one request per home. 10,000 homes means 10,000 page loads, so pacing and proxy budget matter more here than in any other method.

Method 4: Run Zillow's API calls inside a real browser page

Method 4 loads Zillow in Playwright and calls fetch() from inside the page, so every request carries the browser's own cookies, TLS stack, and sensor state.

Best when: Methods 1 to 3 start returning block pages, and you'd rather not reverse-engineer HUMAN's sensor.

To Zillow, the call looks like its own frontend asking for page 2. First, open the page and grab queryState as before.

import json
from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch(headless=False)  # headful: fewer automation tells
    page = browser.new_page(locale="en-US")
    page.goto("https://www.zillow.com/brooklyn-new-york-ny/", wait_until="domcontentloaded")
    page.wait_for_timeout(5000)  # give the sensor script time to run

    raw = page.locator("script#__NEXT_DATA__").text_content()
    query_state = json.loads(raw)["props"]["pageProps"]["searchPageState"]["queryState"]
    # ...the in-page fetch below goes here, inside the same `with` block

Then run the search as JavaScript inside the page. Playwright passes body in and returns the parsed result to Python.

JS_SEARCH = """async (body) => {
  const r = await fetch("/async-create-search-page-state", {
    method: "PUT",
    headers: {"content-type": "application/json"},
    body: JSON.stringify(body),
  });
  return {status: r.status, json: r.ok ? await r.json() : null};
}"""

body = {
    "searchQueryState": {**query_state, "pagination": {"currentPage": 2}},
    "wants": {"cat1": ["listResults"], "cat2": ["total"]},
    "requestId": 3,
}
result = page.evaluate(JS_SEARCH, body)
print(result["status"], len(result["json"]["cat1"]["searchResults"]["listResults"]))

Run it headful. Headless Chromium exposes properties that sensor scripts check for, and on a server xvfb-run python scraper.py gives you a virtual display.

For passive collection, register page.on("response", ...) before page.goto() and keep the search responses as you scroll. See the Playwright network docs.

Tradeoff: each browser needs a few hundred MB of RAM and runs far slower than raw HTTP.

Method 5: Use a stealth browser, then hand its cookies to curl_cffi

Method 5 lets a patched browser earn Zillow's trust once, then reuses those cookies in curl_cffi for fast requests on the same IP.

Best when: you need sustained volume, and running Method 4's browser for every request is too slow.

Stock Playwright Chromium carries automation flags HUMAN looks for. Patchright is a drop-in fork that patches those leaks, and a persistent profile keeps cookies between runs.

from patchright.sync_api import sync_playwright

PROXY = {"server": "http://us.proxy.example:8000",  # sticky US residential IP
         "username": "user-session-zillow1", "password": "pass"}

with sync_playwright() as p:
    ctx = p.chromium.launch_persistent_context(
        user_data_dir="./zillow-profile",  # cookies survive between runs
        channel="chrome", headless=False, no_viewport=True, proxy=PROXY,
    )
    page = ctx.pages[0] if ctx.pages else ctx.new_page()
    page.goto("https://www.zillow.com/brooklyn-new-york-ny/")
    page.wait_for_timeout(8000)  # if Press & Hold appears, do it by hand once
    cookies = ctx.cookies("https://www.zillow.com")
    user_agent = page.evaluate("navigator.userAgent")
    ctx.close()

If Press & Hold appears, complete it by hand once per profile.

Then copy the cookies and the exact User-Agent into a curl_cffi session that goes out through the same proxy.

from curl_cffi import Session

fast = Session(
    impersonate="chrome",
    proxies={"https": "http://user-session-zillow1:[email protected]:8000"},
)
fast.headers.update({"user-agent": user_agent, "accept-language": "en-US,en;q=0.9"})
for c in cookies:
    fast.cookies.set(c["name"], c["value"], domain=c["domain"])

resp = fast.get("https://www.zillow.com/brooklyn-new-york-ny/2_p/", timeout=30)
print(resp.status_code, "__NEXT_DATA__" in resp.text)

Treat those cookies as bound to the IP and fingerprint that earned them. Switch proxies mid-session and they're worthless, so use a sticky US residential or ISP session.

We sell those at Roundproxies, but any provider with sticky US residential sessions will do the job.

Watch the versions: impersonate="chrome" targets one Chrome build. If your installed Chrome is much newer, TLS profile and User-Agent disagree, so keep curl_cffi updated.

Tradeoff: cookies expire. Build a refresh loop: when curl_cffi gets a block page, reopen the browser profile, re-earn the cookies, and resume.

Method 6: Split the map into tiles to get past the 500-result cap

Method 6 splits a search area into smaller map tiles until each tile returns fewer than 500 homes, which gets you every listing in a city.

Best when: you want full coverage of a city or county, beyond the first 500 results.

Zillow's map search returns at most 500 homes per request, however many match. So when a bounding box hits 500, cut it into four quadrants and query each one, recursively.

MAX_PER_QUERY = 500  # Zillow's mapResults ceiling per request

def split(b: dict) -> list:
    """Cut a map bounding box into four equal quadrants."""
    mid_lat = (b["north"] + b["south"]) / 2
    mid_lng = (b["east"] + b["west"]) / 2
    return [
        {"north": b["north"], "south": mid_lat, "west": b["west"], "east": mid_lng},
        {"north": b["north"], "south": mid_lat, "west": mid_lng, "east": b["east"]},
        {"north": mid_lat, "south": b["south"], "west": b["west"], "east": mid_lng},
        {"north": mid_lat, "south": b["south"], "west": mid_lng, "east": b["east"]},
    ]

The crawler below queries a tile and either recurses into its quadrants or stores the results. It reuses put_search() from Method 2.

def crawl_tile(query_state: dict, bounds: dict, seen: dict, depth: int = 0) -> None:
    qs = {**query_state, "mapBounds": bounds, "isMapVisible": True}
    data = put_search({"searchQueryState": qs, "wants": {"cat1": ["mapResults"]},
                       "requestId": depth + 2})
    homes = data["cat1"]["searchResults"]["mapResults"]
    if len(homes) >= MAX_PER_QUERY and depth < 6:
        for child in split(bounds):
            time.sleep(random.uniform(2, 5))
            crawl_tile(query_state, child, seen, depth + 1)
        return
    if len(homes) >= MAX_PER_QUERY:
        print(f"Tile still capped at max depth: {bounds}")  # dense condo block
    for home in homes:
        seen[home.get("zpid") or home.get("detailUrl")] = home  # dedupe by zpid

seed = get_next_data("https://www.zillow.com/brooklyn-new-york-ny/")
qs = seed["props"]["pageProps"]["searchPageState"]["queryState"]
seen = {}
crawl_tile(qs, qs["mapBounds"], seen)
print(f"{len(seen)} unique listings")

Keying the dict on zpid drops duplicates from homes on tile borders.

The depth limit stops one dense condo tower from recursing forever, and the log line shows where coverage is incomplete.

Tradeoff: a dense metro can take hundreds of requests. Pace them, and route the crawl through Method 4 or 5 if plain curl_cffi starts getting blocked.

Which method should you use to scrape Zillow?

Start with Method 2 and add layers only when Zillow pushes back. I'd expect most projects to settle on Methods 2 and 6, with Method 5 held back for recovery.

Your situation Use
Quick look at one search Method 1
All pages of a filtered search (rentals, sold) Method 2
Price history or Zestimate for known homes Method 3
curl_cffi returns 403 even at low volume Method 4
Thousands of requests per day, every day Method 5
Every listing in a city or county Method 6, then Method 3 for details

When you don't need a scraper at all

If you want market trends rather than individual listings, Zillow publishes the data for free. Zillow Research offers CSVs of home values (ZHVI) and rents (ZORI) by metro, city, county, and ZIP.

The files refresh monthly. A CSV download takes seconds and never gets blocked, so save scraping for when you need listing-level rows.

Common errors

Most Zillow scraper failures trace back to four errors, and each one has a specific fix.

"Access to this page has been denied" (HTTP 403)

Why: PerimeterX scored the request as a bot on TLS fingerprint, IP reputation, or missing sensor data.

Fix: update curl_cffi, move to a residential IP, and slow down. Still blocked? Switch to Method 4 or 5.

HTTP 429 Too Many Requests

Why: too many requests from one IP in a short window.

Fix: back off exponentially with random jitter, and spread requests across more IPs. This wrapper handles the retry:

def get_with_backoff(url: str, tries: int = 5):
    for attempt in range(tries):
        resp = session.get(url, timeout=30)
        if resp.status_code != 429:
            return resp
        time.sleep(2 ** attempt + random.uniform(0, 1))  # 1s, 2s, 4s... plus jitter
    raise RuntimeError(f"Still rate-limited after {tries} tries: {url}")

Five tries wait about 31 seconds in total. Our guide to HTTP error 429 covers Retry-After headers too.

KeyError: 'searchPageState' or 'gdpClientCache'

Why: you parsed a block page, or Zillow moved the key.

Fix: print the first 500 characters of the response. Block page? Apply the 403 fix. Real page? Use find_key() to locate the new path.

Search endpoint returns empty listResults

Why: the queryState is stale or malformed, usually from a hand-edited payload.

Fix: re-seed it from a fresh page load as in Method 2, and change filters in the browser, not in code.

A note on responsible use

Zillow's Terms of Use prohibit automated queries, including scraping and bypassing CAPTCHA-style challenges. Read them before you start, and weigh that risk for your use case.

Keep request rates low, skip personal data you don't need, and don't republish listing photos; photographers and brokers own them. None of this is legal advice.

FAQ

It's a gray area. US courts have mostly not treated public-page scraping as CFAA hacking (hiQ v. LinkedIn), but Zillow's Terms of Use ban it. Ask a lawyer before commercial use.

Does Zillow have a public API?

Not for general use. Zillow retired its public API in 2021, and listing data now goes to approved partners through Bridge Interactive. For market-level statistics, Zillow Research publishes free CSV downloads.

Do you need proxies to scrape Zillow?

Yes, once you go past a few hundred requests a day. Use residential or ISP proxies with sticky sessions, because datacenter IPs get flagged quickly and rotating mid-session invalidates your cookies.

How do I scrape Zillow rentals or sold homes?

Seed Method 2 with a filtered URL. Apply the For Rent or Sold filter in your browser, copy the URL, and the scraper inherits the filter through the page's queryState.

How many requests can I send before Zillow blocks me?

Zillow doesn't publish a limit, and the threshold moves with your IP and fingerprint. Start at one request every 3 to 7 seconds per IP, then adjust to your block rate.

Wrapping up

Zillow gets much easier once you stop fighting the HTML: the data already ships as JSON, and the search endpoint hands you pagination and filters.

Start with Methods 1 and 2 on curl_cffi, add tiling for full coverage, and keep a stealth browser profile ready for when PerimeterX tightens up.

Building a multi-country property dataset? Our ImmobilienScout24 scraping guide covers Germany's largest real estate portal.