Knowledgebase

Python requests tutorial: tested code for 2026

requests.get(url) is the first line most people write when Python needs to talk to the web.

It's also where most scripts pick up their first production bug: no timeout, so one unresponsive server freezes the process until someone notices.

This guide covers how to use requests in Python from the first GET call to a session you can leave running unattended.

You'll send parameters, JSON, files, and headers, then add timeouts, retries, proxies, streaming, and threads.

I ran every snippet here against requests 2.33.1 and urllib3 2.6.3 on Python 3.12, using a local test server.

That testing turned up five popular snippets that fail without raising anything, four of which lived in the previous version of this page. They get their own section.

What is the Python requests library?

Python requests is a third-party HTTP library built on urllib3. You call requests.get() or requests.post() and get back a Response object holding the status code, headers, and decoded body. Install it with pip install requests, use it for APIs, webhooks, and static pages, and pass a timeout on every call.

Under the hood, each call builds a PreparedRequest and hands it to an HTTPAdapter. The adapter borrows a socket from a urllib3 connection pool and sends the bytes.

That chain matters later. Timeouts, retries, and pool sizes all live on the adapter, which is why a few popular "session-wide" settings do nothing.

The current release is 2.34.2 (May 2026), and it supports Python 3.10 and newer. It's synchronous, speaks HTTP/1.1 only, and never runs JavaScript.

If one of those limits rules it out, pick from this table instead:

Library Use it when Async HTTP/2
requests Scripts, API clients, static-page scraping No No
httpx You want a requests-style API plus async or HTTP/2 Yes Yes, via httpx[http2]
aiohttp Hundreds of concurrent connections inside an asyncio app Yes No
urllib.request You can't install third-party packages No No

For most scripts, requests is still the right default. The API is smaller, the docs are better, and every Stack Overflow answer assumes it.

How to install requests

Install it into a virtual environment so each project pins its own version. The last line confirms which release landed.

python -m venv .venv
source .venv/bin/activate        # Windows: .venv\Scripts\activate
python -m pip install requests
python -c "import requests; print(requests.__version__)"

Use python -m pip instead of bare pip. On machines with several Pythons, pip can install into an interpreter your script never runs.

The install also pulls in urllib3 (connections), certifi (the CA bundle for TLS checks), charset_normalizer (encoding detection), and idna (international domains). Remember certifi: it's behind most certificate errors later.

Make your first GET request

This call fetches a public GitHub profile and reads three things off the response.

import requests

response = requests.get("https://api.github.com/users/github", timeout=10)

print(response.status_code)              # 200
print(response.headers["content-type"])  # application/json; charset=utf-8
print(response.json()["name"])           # GitHub

Header lookups are case-insensitive, so "Content-Type" and "content-type" return the same value. .json() parses the body into a dict or list.

What the Response object gives you

You'll use most of these attributes within your first week:

Attribute Returns Notes
.status_code int 200, 404, 503...
.ok bool True for any status below 400
.text str Body decoded with .encoding
.content bytes Raw body; use for images, PDFs, zips
.json() dict / list Raises JSONDecodeError on non-JSON
.headers dict-like Case-insensitive keys
.url str Final URL after redirects
.history list Redirect responses, oldest first
.elapsed timedelta Time until headers arrived, not the full body
.request PreparedRequest What was sent: URL, headers, body

A response in a boolean context follows .ok, so if response: is False for a 503. That's handy, and easy to misread as "the request happened".

One encoding trap deserves a mention. If a text/html response has no charset in its Content-Type, requests falls back to ISO-8859-1, and accented characters come out garbled.

Fix it by setting response.encoding = response.apparent_encoding before reading .text. That swaps in charset_normalizer's guess.

Send query parameters

Pass a dict to params and let requests handle encoding. This example uses httpbin, which echoes back what it received.

import requests

params = {
    "q": "python requests",
    "page": 2,
    "debug": None,                # None values are dropped from the URL
    "tag": ["http", "scraping"],  # lists repeat the key
}
response = requests.get("https://httpbin.org/get", params=params, timeout=10)
print(response.url)
# https://httpbin.org/get?q=python+requests&page=2&tag=http&tag=scraping

Spaces become +, and None disappears entirely. That second behavior is useful for optional filters: build one dict and let unset values vanish.

Send data with POST, PUT, PATCH, and DELETE

The keyword you pick decides the body format and the Content-Type header:

Argument Body format Content-Type set for you
json= JSON application/json
data= (dict) Form-encoded application/x-www-form-urlencoded
files= Multipart multipart/form-data; boundary=...
data= (str or bytes) Sent as-is None; set it yourself

Most modern APIs want JSON. Older ones and HTML forms want form encoding.

import requests

# JSON body: serialized for you
r = requests.post("https://httpbin.org/post", json={"name": "Ada", "role": "admin"}, timeout=10)
print(r.json()["json"])   # {'name': 'Ada', 'role': 'admin'}

# Form body
r = requests.post("https://httpbin.org/post", data={"name": "Ada"}, timeout=10)
print(r.json()["form"])   # {'name': 'Ada'}

Don't pass data= and json= together. When I tried it, requests sent the form body and dropped the JSON without a warning.

Upload a file

Open files in binary mode and pass a tuple of filename, file object, and MIME type.

import requests

with open("report.csv", "rb") as fh:
    files = {"file": ("report.csv", fh, "text/csv")}
    r = requests.post(
        "https://httpbin.org/post",
        files=files,
        data={"note": "weekly export"},  # extra form fields ride along
        timeout=30,
    )
print(list(r.json()["files"]))  # ['file']

Never set Content-Type by hand for uploads. Requests generates a multipart boundary and writes it into that header; overwrite it and the server can't parse the body.

PUT, PATCH, and DELETE

These share the same signature. PUT replaces a whole resource, PATCH changes some fields, DELETE removes it.

import requests

base = "https://httpbin.org"
requests.put(f"{base}/put", json={"id": 7, "name": "Ada", "role": "admin"}, timeout=10)
requests.patch(f"{base}/patch", json={"role": "viewer"}, timeout=10)

r = requests.delete(f"{base}/delete", timeout=10)
print(r.status_code)  # 200 here; many real APIs return 204

A 204 response has no body, so calling .json() on it raises. Check the status code before parsing.

Set headers and a User-Agent

First, see what requests sends when you set nothing. My local test server received these headers from a bare call, and httpbin shows the same set:

import requests

print(requests.get("https://httpbin.org/headers", timeout=10).json()["headers"])
# {'User-Agent': 'python-requests/2.33.1', 'Accept-Encoding': 'gzip, deflate',
#  'Accept': '*/*', 'Connection': 'keep-alive', ...}

Accept-Encoding gains br or zstd if those packages are installed. The part that matters is the User-Agent: it announces a script to every server you touch.

For APIs, that's fine. Replace it with something that identifies you, like "inventory-sync/2.0 ([email protected])". Operators can contact you instead of blocking you.

For HTML pages, send what a browser sends:

import requests

headers = {
    # Copy the current string from your own browser's devtools
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
                  "(KHTML, like Gecko) Chrome/140.0.0.0 Safari/537.36",
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
    "Accept-Language": "en-US,en;q=0.9",
}
r = requests.get("https://example.com", headers=headers, timeout=10)
print(r.status_code)

This gets past the laziest filters. For anything tougher, see how user-agent rotation works and why one UA across thousands of requests stands out.

Headers have a ceiling, though. Anti-bot vendors also read the TLS handshake, and requests produces a Python-shaped one you can't change. More on that in how TLS fingerprinting works.

Handle cookies and redirects

A Session stores cookies the server sets and sends them back automatically. Here httpbin sets a cookie, then reads it back:

import requests

with requests.Session() as s:
    s.get("https://httpbin.org/cookies/set?session_id=abc123", timeout=10)
    print(s.cookies.get("session_id"))  # abc123

    r = s.get("https://httpbin.org/cookies", timeout=10)
    print(r.json())                     # {'cookies': {'session_id': 'abc123'}}

Plain requests.get() calls start with an empty jar every time. Any login flow needs a session.

Redirects are followed by default, up to 30 hops. Use allow_redirects=False when you need the Location header itself:

import requests

r = requests.get("http://github.com", timeout=10)
print(r.url, [h.status_code for h in r.history])  # https://github.com/ [301]

r = requests.get("http://github.com", allow_redirects=False, timeout=10)
print(r.status_code, r.headers["Location"])       # 301 https://github.com/

When a redirect points to a different host, requests strips your Authorization header. That's a security feature, and it explains the classic "works on the first URL, 401 after the redirect" bug.

Authenticate your requests

HTTP Basic auth takes a tuple. Requests base64-encodes it into the header for you.

import requests

r = requests.get("https://httpbin.org/basic-auth/user/pass", auth=("user", "pass"), timeout=10)
print(r.status_code)  # 200

Basic auth is only base64, not encryption. Send it over HTTPS or not at all.

For bearer tokens, a small AuthBase subclass beats a hand-built header. It attaches cleanly to a session and keeps the token format in one place.

import os
import requests
from requests.auth import AuthBase

class BearerAuth(AuthBase):
    def __init__(self, token):
        self.token = token

    def __call__(self, request):
        request.headers["Authorization"] = f"Bearer {self.token}"
        return request

r = requests.get("https://httpbin.org/bearer", auth=BearerAuth(os.environ["API_TOKEN"]), timeout=10)
print(r.json())  # {'authenticated': True, 'token': '...'}

There's a second reason to prefer auth=. If your ~/.netrc has an entry for the host, requests applies it and overwrites an Authorization header you set by hand.

Keep tokens in environment variables, as above, and out of your repo. For OAuth 1 and 2 flows, the open-source requests-oauthlib package plugs into the same auth= argument.

Set timeouts and handle errors

Requests has no default timeout. Without one, a server that accepts the connection and never answers holds your script indefinitely.

import requests

requests.get("https://api.github.com", timeout=10)       # 10s for connect and 10s for read
requests.get("https://api.github.com", timeout=(3, 30))  # 3s to connect, 30s between bytes

The tuple form is worth the extra characters. A dead host should fail fast, while a slow report endpoint may need longer to send its first byte.

The read timeout is the gap allowed between bytes, not a limit on the whole download. A server that trickles one byte every nine seconds never trips timeout=10.

When you need a hard deadline, stream the body and check the clock yourself:

import time
import requests

def get_with_deadline(url, deadline=30):
    start = time.monotonic()
    chunks = []
    with requests.get(url, stream=True, timeout=(5, 10)) as r:
        r.raise_for_status()
        for chunk in r.iter_content(64 * 1024):
            if time.monotonic() - start > deadline:
                raise TimeoutError(f"{url} took longer than {deadline}s")
            chunks.append(chunk)
    return b"".join(chunks)

The check only runs when a chunk arrives, so worst-case overrun is one read timeout. That's usually close enough.

Catch the right exceptions

Every requests exception inherits from RequestException. The subclasses tell you whether retrying makes sense:

Exception Raised when Retry it?
ConnectTimeout TCP or TLS setup ran past the connect timeout Yes. The server never saw the request
ReadTimeout The server accepted, then stalled Only for idempotent calls
ConnectionError DNS failure, refused or reset connection Usually
HTTPError raise_for_status() on a 4xx or 5xx 429 and 5xx yes, other 4xx no
JSONDecodeError .json() on a non-JSON body No. Read the body first
MissingSchema, InvalidURL The URL is malformed No
TooManyRedirects More than 30 redirects No

ConnectTimeout inherits from both ConnectionError and Timeout. Put it first in your except chain, or the ConnectionError branch swallows it.

import requests
from requests import exceptions as rex

def fetch_json(url):
    try:
        r = requests.get(url, timeout=(5, 15))
        r.raise_for_status()
        return r.json()
    except rex.ConnectTimeout:
        print("Couldn't connect in time")
    except rex.ReadTimeout:
        print("Server accepted the connection, then stalled")
    except rex.ConnectionError:
        print("DNS failure, refused, or reset")
    except rex.HTTPError as e:
        print(f"Bad status: {e.response.status_code}")
    except rex.JSONDecodeError:
        print(f"Expected JSON, got: {r.text[:200]!r}")
    except rex.RequestException as e:
        print(f"Request failed: {e}")

The JSON branch prints the first 200 characters of the body. When an API starts returning an HTML error page, that one line saves you twenty minutes.

Reuse connections with a Session

Every bare requests.get() opens a fresh TCP connection and, for HTTPS, a fresh TLS handshake. A Session keeps finished connections in a pool and reuses them for the same host.

The saving is one or two network round trips per request. Against a server 100 ms away, that's 200 ms or more on every call after the first.

Sessions also carry headers, cookies, and auth across calls:

import requests

repos = ["psf/requests", "encode/httpx", "aio-libs/aiohttp"]

with requests.Session() as session:
    session.headers.update({"Accept": "application/vnd.github+json", "User-Agent": "repo-audit/1.0"})
    for repo in repos:
        r = session.get(f"https://api.github.com/repos/{repo}", timeout=10)
        r.raise_for_status()
        print(repo, r.json()["stargazers_count"])

The with block closes pooled sockets on exit. In a long-running service, create one session at startup and reuse it for the life of the process.

One trap: session.timeout = 5 looks right and does nothing, because Session has no timeout setting.

Python quietly creates a new attribute that requests never reads. When I pointed a session with that line at a 3-second endpoint, it waited the full 3 seconds.

The fix is an adapter, covered in the production session section.

Retry failed requests automatically

Requests doesn't retry failed HTTP responses on its own. You attach a urllib3 Retry policy to an adapter, and the adapter handles the loop.

import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

retry = Retry(
    total=3,
    backoff_factor=1,                            # sleeps 0s, 2s, 4s between attempts
    status_forcelist=[429, 500, 502, 503, 504],
)
adapter = HTTPAdapter(max_retries=retry)

session = requests.Session()
session.mount("https://", adapter)
session.mount("http://", adapter)
r = session.get("https://httpbin.org/status/200", timeout=10)

The first retry fires immediately.

Later ones sleep backoff_factor * 2 ** (n - 1) seconds, capped at 120 by default. I timed three failed retries against an always-503 endpoint: 6.05 seconds total.

Retry-After is honored by default on 413, 429, and 503 responses. If a rate-limited API tells you to wait 30 seconds, urllib3 waits 30 seconds.

By default, only idempotent methods retry: GET, HEAD, PUT, DELETE, OPTIONS, and TRACE. POST is excluded on purpose.

Adding "POST" to allowed_methods can charge a card twice when the first attempt succeeded but the response got lost. Only do it if the API accepts idempotency keys.

When retries run out on a status code, you get RetryError, not the last response. Pass raise_on_status=False to Retry if you'd rather inspect the final 503 yourself.

Retries also change which exception reaches you. Once the policy gives up on read timeouts, requests raises ConnectionError with Caused by ReadTimeoutError in the message.

I hit this while testing: an except ReadTimeout branch never fired against a session with retries mounted. Catch ConnectionError there, or check isinstance(e.args[0].reason, ReadTimeoutError) if you need to tell them apart.

For deeper coverage of rate limits specifically, see how to handle HTTP error 429.

Use proxies with Python requests

Pass a dict that maps URL schemes to proxy URLs. Credentials go inside the URL.

import os
import requests

proxy = os.environ["PROXY_URL"]  # e.g. http://user:[email protected]:8000
proxies = {"http": proxy, "https": proxy}

r = requests.get("https://httpbin.org/ip", proxies=proxies, timeout=15)
print(r.json())  # {'origin': '<the proxy exit IP>'}

The "https" key picks the proxy for HTTPS URLs. Its value still usually starts with http://, because requests opens a CONNECT tunnel through it.

Residential and datacenter endpoints from Roundproxies use the same user:pass@host:port format, so this dict works unchanged.

A few details catch people out:

  • SOCKS needs an extra. Run pip install "requests[socks]", then use socks5h:// rather than socks5://. The h resolves DNS at the proxy, so lookups don't leak from your machine.
  • Environment variables count. Requests reads HTTP_PROXY, HTTPS_PROXY, and NO_PROXY. On CI runners or corporate laptops, those can route traffic you didn't expect.
  • session.proxies can lose. The requests docs warn that environment proxies override it. Pass proxies= on each call, or set session.trust_env = False.

Rotation, sticky sessions, and failover get the full treatment in the Python requests proxy guide.

Stream large downloads

By default, requests loads the entire body into memory. A 4 GB file means 4 GB of RAM. Add stream=True and read in chunks:

import requests

def download(url, path, chunk_size=256 * 1024):
    with requests.get(url, stream=True, timeout=(5, 30)) as r:
        r.raise_for_status()
        total = int(r.headers.get("Content-Length", 0))
        with open(path, "wb") as fh:
            for chunk in r.iter_content(chunk_size):
                fh.write(chunk)
                if total:
                    # raw.tell() counts bytes off the wire, same unit as Content-Length
                    print(f"\r{r.raw.tell() / total:.0%}", end="")
    print()

download("https://example.com/export.zip", "export.zip")

The progress math uses r.raw.tell() for a reason. iter_content hands you decompressed bytes, while Content-Length counts compressed bytes. Summing len(chunk) is the bug in most tutorials; details in the next section.

Chunked responses have no Content-Length, so the if total guard skips the percentage instead of dividing by zero.

For line-delimited feeds such as NDJSON exports, use iter_lines():

import json
import requests

with requests.get("https://example.com/events.ndjson", stream=True, timeout=(5, 60)) as r:
    for line in r.iter_lines():
        if line:  # skip keep-alive blank lines
            print(json.loads(line))

Always close streamed responses, which the with block does. An unread streamed response keeps its connection checked out of the pool.

Make concurrent requests with threads

Requests blocks while it waits, but it releases the GIL during network I/O. A thread pool gets you real parallelism for fetching.

from concurrent.futures import ThreadPoolExecutor
import requests
from requests.adapters import HTTPAdapter

WORKERS = 20
session = requests.Session()
# pool_maxsize must be >= WORKERS, or urllib3 throws finished sockets away
session.mount("https://", HTTPAdapter(pool_connections=4, pool_maxsize=WORKERS))

def fetch(url):
    r = session.get(url, timeout=10)
    return r.status_code, url

urls = [f"https://httpbin.org/anything/{i}" for i in range(100)]
with ThreadPoolExecutor(max_workers=WORKERS) as pool:
    for status, url in pool.map(fetch, urls):
        print(status, url)

The default pool holds 10 connections per host. When I ran 20 threads against a default session, urllib3 logged this 10 times:

Connection pool is full, discarding connection: 127.0.0.1. Connection pool size: 10

Nothing crashes, but every discarded socket means a new handshake later. Setting pool_maxsize=20 brought the count to zero.

Sharing one session across threads works for stateless GETs, but requests doesn't promise it's thread-safe.

If workers log in or depend on their own cookies, give each thread its own session through threading.local().

Past a few hundred concurrent requests, threads get expensive. That's the point to switch to httpx's async client or aiohttp.

5 requests snippets that quietly misbehave

These snippets show up across tutorials, Stack Overflow answers, and, until this rewrite, this page.

None of them raises an error. I ran each against a local test server with requests 2.33.1 and urllib3 2.6.3.

Snippet What it claims What I measured Fix
Retry(total=3, backoff_factor=1) Waits 1s, 2s, 4s Waits 0s, 2s, 4s. Three retries finished in 6.05s Expect an immediate first retry
Progress bar summing len(chunk) Percent downloaded 86,957% on a gzip response: 230 bytes on the wire, 200,000 decoded Divide r.raw.tell() by Content-Length
TimeoutHTTPAdapter that sets kwargs["timeout"] A default timeout Replaced my explicit timeout=10 with its 1s default Only fill in when timeout is None
session.timeout = 5 Session-wide timeout Ignored. A 3s endpoint took the full 3.0s Use an adapter default
int(r.headers["Retry-After"]) Seconds to wait ValueError when the server sends an HTTP date Parse both formats

The last one crashes your error handler in the one moment you need it. Retry-After can legally be seconds or a date like Wed, 21 Oct 2037 07:28:00 GMT.

This helper handles both and falls back to a default:

from datetime import datetime, timezone
from email.utils import parsedate_to_datetime

def retry_after_seconds(value, default=60.0):
    if value is None:
        return default
    value = value.strip()
    if value.isdigit():
        return float(value)
    try:
        when = parsedate_to_datetime(value)
    except (TypeError, ValueError):
        return default
    return max(0.0, (when - datetime.now(timezone.utc)).total_seconds())

Call it with r.headers.get("Retry-After"). A missing header returns the default instead of raising KeyError.

Build a production-ready session

Everything above fits into two short pieces you can drop into any project. First, an adapter that applies a default timeout only when the caller didn't pass one:

from requests.adapters import HTTPAdapter

class DefaultTimeoutAdapter(HTTPAdapter):
    """HTTPAdapter that fills in a timeout only when the caller omitted one."""

    def __init__(self, *args, timeout=(5, 30), **kwargs):
        self.timeout = timeout
        super().__init__(*args, **kwargs)

    def send(self, request, **kwargs):
        if kwargs.get("timeout") is None:
            kwargs["timeout"] = self.timeout
        return super().send(request, **kwargs)

Session always passes a timeout key to the adapter, set to None when you didn't specify one. Checking for None is what keeps explicit values intact.

Then a factory that wires retries, jitter, pool size, and an honest User-Agent into one session:

import requests
from urllib3.util.retry import Retry

def make_session(user_agent, pool_size=10, retries=3):
    retry = Retry(
        total=retries,
        backoff_factor=0.5,
        backoff_jitter=0.3,  # spreads retries out when many workers fail together
        status_forcelist=[429, 500, 502, 503, 504],
    )
    adapter = DefaultTimeoutAdapter(max_retries=retry, pool_maxsize=pool_size)
    session = requests.Session()
    session.headers["User-Agent"] = user_agent
    session.mount("https://", adapter)
    session.mount("http://", adapter)
    return session

session = make_session("inventory-sync/2.0 ([email protected])", pool_size=20)
r = session.get("https://httpbin.org/get")  # 5s connect, 30s read, 3 retries

backoff_jitter needs urllib3 2.0 or newer, which any current requests install already has. Without jitter, fifty workers that fail together retry together and hit the server in lockstep.

Build the session once, pass it around, and every call gets sane defaults without anyone remembering them.

Troubleshooting common Python requests errors

These are the exact messages from my test runs, so you can match them against your traceback.

ModuleNotFoundError: No module named 'requests'

Why: The package went into a different interpreter than the one running your script.

Fix: Print sys.executable inside the failing script, then run <that path> -m pip install requests.

MissingSchema: Invalid URL 'example.com': No scheme supplied. Perhaps you meant https://example.com?

Why: Requests needs http:// or https:// in front of every URL.

Fix: Add the scheme. When URLs come from user input or a CSV, normalize them before the request.

ConnectionError: ... Max retries exceeded with url: / (Caused by NewConnectionError(... [Errno 111] Connection refused))

Why: Nothing is listening on that host and port. "Max retries exceeded" is urllib3's wording and appears even when you configured zero retries.

Fix: Read the Caused by part. It tells you whether DNS failed, the port refused, or the connection reset. Check the host, port, and any proxy settings.

ReadTimeout: HTTPConnectionPool(...): Read timed out. (read timeout=1)

Why: The connection opened, but the server went quiet for longer than your read timeout.

Fix: Raise the second value of the timeout tuple for slow endpoints. Retry only if the method is idempotent.

JSONDecodeError: Expecting value: line 1 column 1 (char 0)

Why: The body isn't JSON. Usually it's an HTML error page, a CAPTCHA, or an empty 204.

Fix: Print r.status_code and r.text[:200] before calling .json(). The body tells you which of those it is.

SSLError: [SSL: CERTIFICATE_VERIFY_FAILED] certificate verify failed: unable to get local issuer certificate

Why: The server's certificate chain doesn't lead to a CA in certifi's bundle. Corporate TLS inspection and self-signed internal services are the usual causes.

Fix: Run python -m pip install --upgrade certifi. For an internal CA, pass verify="/path/to/company-ca.pem". Keep verify=False out of production.

RetryError: ... (Caused by ResponseError('too many 503 error responses'))

Why: Your Retry policy ran out of attempts on a status in status_forcelist.

Fix: Catch requests.exceptions.RetryError, or set raise_on_status=False to get the final response back.

Connection pool is full, discarding connection

Why: More threads than pool_maxsize are talking to one host.

Fix: Mount an HTTPAdapter with pool_maxsize at least equal to your worker count.

403 Forbidden on a page that loads in your browser

Why: The server flagged the request as automated. The python-requests User-Agent is the first suspect; the TLS fingerprint is the second.

Fix: Send browser headers as shown earlier. If a 403 survives that, the block is below the HTTP layer, and requests alone won't clear it.

When requests is the wrong tool

Requests covers most HTTP work, but four situations call for something else:

  • The data appears only after JavaScript runs. Requests fetches the HTML the server sends. Use browser automation, or find the JSON endpoint the page calls. The Python web scraping guide covers both routes.
  • You need hundreds of requests in flight. Use httpx's AsyncClient or aiohttp.
  • The server wants HTTP/2. Use httpx with the http2 extra.
  • You're blocked on the TLS fingerprint. The open-source curl_cffi library impersonates browser handshakes behind a requests-style API.

For everything else, requests plus the session factory above is enough.

FAQ

Is requests part of the Python standard library?

No. Requests is a third-party package you install with pip install requests. The standard library's urllib.request handles HTTP without installs, but you encode parameters and parse responses by hand.

What is the default timeout in Python requests?

There isn't one. Requests waits indefinitely for a response unless you pass timeout=.

Connection attempts eventually fail at the operating system level (roughly two minutes on Linux), but a stalled read can hang forever.

How do I send JSON with Python requests?

Pass a dict to the json= argument: requests.post(url, json={"key": "value"}, timeout=10). Requests serializes it and sets Content-Type: application/json for you.

What's the difference between data and json in requests?

data= sends a dict as form-encoded fields, the format HTML forms use.

json= sends it as a JSON document. Pass only one of them; with both, requests sends the form data and silently drops the JSON.

Is requests.Session thread-safe?

Not officially. Sharing one session for stateless GET requests works in practice if pool_maxsize matches your thread count. For logins or per-worker cookies, create one session per thread.

Does Python requests support async?

No. Requests is synchronous and each call blocks its thread. For async code, httpx offers a nearly identical API through httpx.AsyncClient, and aiohttp suits very high concurrency.

Wrapping up

The mental model is short: requests.get() for one-offs, a Session for anything repeated, and an adapter for timeouts, retries, and pool size.

Start by copying the make_session() factory into your project and routing every call through it.

It closes the timeout gap, retries the failures worth retrying, and sends a User-Agent you'd be comfortable explaining.

When your target starts blocking you, head to the proxy guide for Python requests. The official requests documentation and the urllib3 Retry reference cover every argument this guide skipped.