> ## Content Index
> Fetch the complete content index at: https://roundproxies.com/blog/llms.txt
> Use this file to discover other available public pages before exploring further.

# Python requests tutorial: tested code for 2026
- URL: https://roundproxies.com/blog/requests-python/
- Published: 2025-10-06T22:16:48.000Z
- Updated: 2026-09-24T13:35:08.000Z
- Description: How to use Python requests: GET, POST, sessions, retries, proxies. Every snippet tested.
- Author: Marius Bernard
- Tags: Knowledgebase, #dated-68e43b524fa498a2c6ee1f5c

`requests.get(url)` is the first line most people write when Python needs to talk to the web.

It's also where most scripts pick up their first production bug: no timeout, so one unresponsive server freezes the process until someone notices.

This guide covers how to use requests in Python from the first GET call to a session you can leave running unattended.

You'll send parameters, JSON, files, and headers, then add timeouts, retries, proxies, streaming, and threads.

I ran every snippet here against requests 2.33.1 and urllib3 2.6.3 on Python 3.12, using a local test server.

That testing turned up five popular snippets that fail without raising anything, four of which lived in the previous version of this page. They get their own section.

## What is the Python requests library?

Python requests is a third-party HTTP library built on urllib3\. You call `requests.get()` or `requests.post()` and get back a Response object holding the status code, headers, and decoded body. Install it with `pip install requests`, use it for APIs, webhooks, and static pages, and pass a `timeout` on every call.

Under the hood, each call builds a `PreparedRequest` and hands it to an `HTTPAdapter`. The adapter borrows a socket from a urllib3 connection pool and sends the bytes.

That chain matters later. Timeouts, retries, and pool sizes all live on the adapter, which is why a few popular "session-wide" settings do nothing.

The current release is 2.34.2 (May 2026), and it supports Python 3.10 and newer. It's synchronous, speaks HTTP/1.1 only, and never runs JavaScript.

If one of those limits rules it out, pick from this table instead:

| Library        | Use it when                                              | Async | HTTP/2                  |
| -------------- | -------------------------------------------------------- | ----- | ----------------------- |
| requests       | Scripts, API clients, static-page scraping               | No    | No                      |
| httpx          | You want a requests-style API plus async or HTTP/2       | Yes   | Yes, via httpx\[http2\] |
| aiohttp        | Hundreds of concurrent connections inside an asyncio app | Yes   | No                      |
| urllib.request | You can't install third-party packages                   | No    | No                      |

For most scripts, requests is still the right default. The API is smaller, the docs are better, and every Stack Overflow answer assumes it.

## How to install requests

Install it into a virtual environment so each project pins its own version. The last line confirms which release landed.

```bash
python -m venv .venv
source .venv/bin/activate        # Windows: .venv\Scripts\activate
python -m pip install requests
python -c "import requests; print(requests.__version__)"

```

Use `python -m pip` instead of bare `pip`. On machines with several Pythons, `pip` can install into an interpreter your script never runs.

The install also pulls in urllib3 (connections), certifi (the CA bundle for TLS checks), charset\_normalizer (encoding detection), and idna (international domains). Remember certifi: it's behind most certificate errors later.

## Make your first GET request

This call fetches a public GitHub profile and reads three things off the response.

```python
import requests

response = requests.get("https://api.github.com/users/github", timeout=10)

print(response.status_code)              # 200
print(response.headers["content-type"])  # application/json; charset=utf-8
print(response.json()["name"])           # GitHub

```

Header lookups are case-insensitive, so `"Content-Type"` and `"content-type"` return the same value. `.json()` parses the body into a dict or list.

### What the Response object gives you

You'll use most of these attributes within your first week:

| Attribute     | Returns         | Notes                                         |
| ------------- | --------------- | --------------------------------------------- |
| .status\_code | int             | 200, 404, 503...                              |
| .ok           | bool            | True for any status below 400                 |
| .text         | str             | Body decoded with .encoding                   |
| .content      | bytes           | Raw body; use for images, PDFs, zips          |
| .json()       | dict / list     | Raises JSONDecodeError on non-JSON            |
| .headers      | dict-like       | Case-insensitive keys                         |
| .url          | str             | Final URL after redirects                     |
| .history      | list            | Redirect responses, oldest first              |
| .elapsed      | timedelta       | Time until headers arrived, not the full body |
| .request      | PreparedRequest | What was sent: URL, headers, body             |

A response in a boolean context follows `.ok`, so `if response:` is `False` for a 503\. That's handy, and easy to misread as "the request happened".

One encoding trap deserves a mention. If a `text/html` response has no `charset` in its Content-Type, requests falls back to ISO-8859-1, and accented characters come out garbled.

Fix it by setting `response.encoding = response.apparent_encoding` before reading `.text`. That swaps in charset\_normalizer's guess.

## Send query parameters

Pass a dict to `params` and let requests handle encoding. This example uses httpbin, which echoes back what it received.

```python
import requests

params = {
    "q": "python requests",
    "page": 2,
    "debug": None,                # None values are dropped from the URL
    "tag": ["http", "scraping"],  # lists repeat the key
}
response = requests.get("https://httpbin.org/get", params=params, timeout=10)
print(response.url)
# https://httpbin.org/get?q=python+requests&page=2&tag=http&tag=scraping

```

Spaces become `+`, and `None` disappears entirely. That second behavior is useful for optional filters: build one dict and let unset values vanish.

## Send data with POST, PUT, PATCH, and DELETE

The keyword you pick decides the body format and the Content-Type header:

| Argument             | Body format  | Content-Type set for you          |
| -------------------- | ------------ | --------------------------------- |
| json=                | JSON         | application/json                  |
| data= (dict)         | Form-encoded | application/x-www-form-urlencoded |
| files=               | Multipart    | multipart/form-data; boundary=... |
| data= (str or bytes) | Sent as-is   | None; set it yourself             |

Most modern APIs want JSON. Older ones and HTML forms want form encoding.

```python
import requests

# JSON body: serialized for you
r = requests.post("https://httpbin.org/post", json={"name": "Ada", "role": "admin"}, timeout=10)
print(r.json()["json"])   # {'name': 'Ada', 'role': 'admin'}

# Form body
r = requests.post("https://httpbin.org/post", data={"name": "Ada"}, timeout=10)
print(r.json()["form"])   # {'name': 'Ada'}

```

Don't pass `data=` and `json=` together. When I tried it, requests sent the form body and dropped the JSON without a warning.

### Upload a file

Open files in binary mode and pass a tuple of filename, file object, and MIME type.

```python
import requests

with open("report.csv", "rb") as fh:
    files = {"file": ("report.csv", fh, "text/csv")}
    r = requests.post(
        "https://httpbin.org/post",
        files=files,
        data={"note": "weekly export"},  # extra form fields ride along
        timeout=30,
    )
print(list(r.json()["files"]))  # ['file']

```

Never set `Content-Type` by hand for uploads. Requests generates a multipart boundary and writes it into that header; overwrite it and the server can't parse the body.

### PUT, PATCH, and DELETE

These share the same signature. PUT replaces a whole resource, PATCH changes some fields, DELETE removes it.

```python
import requests

base = "https://httpbin.org"
requests.put(f"{base}/put", json={"id": 7, "name": "Ada", "role": "admin"}, timeout=10)
requests.patch(f"{base}/patch", json={"role": "viewer"}, timeout=10)

r = requests.delete(f"{base}/delete", timeout=10)
print(r.status_code)  # 200 here; many real APIs return 204

```

A 204 response has no body, so calling `.json()` on it raises. Check the status code before parsing.

## Set headers and a User-Agent

First, see what requests sends when you set nothing. My local test server received these headers from a bare call, and httpbin shows the same set:

```python
import requests

print(requests.get("https://httpbin.org/headers", timeout=10).json()["headers"])
# {'User-Agent': 'python-requests/2.33.1', 'Accept-Encoding': 'gzip, deflate',
#  'Accept': '*/*', 'Connection': 'keep-alive', ...}

```

`Accept-Encoding` gains `br` or `zstd` if those packages are installed. The part that matters is the User-Agent: it announces a script to every server you touch.

For APIs, that's fine. Replace it with something that identifies you, like `"inventory-sync/2.0 (+ops@example.com)"`. Operators can contact you instead of blocking you.

For HTML pages, send what a browser sends:

```python
import requests

headers = {
    # Copy the current string from your own browser's devtools
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
                  "(KHTML, like Gecko) Chrome/140.0.0.0 Safari/537.36",
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
    "Accept-Language": "en-US,en;q=0.9",
}
r = requests.get("https://example.com", headers=headers, timeout=10)
print(r.status_code)

```

This gets past the laziest filters. For anything tougher, see [how user-agent rotation works](https://roundproxies.com/blog/user-agent-rotation/) and why one UA across thousands of requests stands out.

Headers have a ceiling, though. Anti-bot vendors also read the TLS handshake, and requests produces a Python-shaped one you can't change. More on that in [how TLS fingerprinting works](https://roundproxies.com/blog/what-is-tls-fingerprint/).

## Handle cookies and redirects

A `Session` stores cookies the server sets and sends them back automatically. Here httpbin sets a cookie, then reads it back:

```python
import requests

with requests.Session() as s:
    s.get("https://httpbin.org/cookies/set?session_id=abc123", timeout=10)
    print(s.cookies.get("session_id"))  # abc123

    r = s.get("https://httpbin.org/cookies", timeout=10)
    print(r.json())                     # {'cookies': {'session_id': 'abc123'}}

```

Plain `requests.get()` calls start with an empty jar every time. Any login flow needs a session.

Redirects are followed by default, up to 30 hops. Use `allow_redirects=False` when you need the `Location` header itself:

```python
import requests

r = requests.get("http://github.com", timeout=10)
print(r.url, [h.status_code for h in r.history])  # https://github.com/ [301]

r = requests.get("http://github.com", allow_redirects=False, timeout=10)
print(r.status_code, r.headers["Location"])       # 301 https://github.com/

```

When a redirect points to a different host, requests strips your `Authorization` header. That's a security feature, and it explains the classic "works on the first URL, 401 after the redirect" bug.

## Authenticate your requests

HTTP Basic auth takes a tuple. Requests base64-encodes it into the header for you.

```python
import requests

r = requests.get("https://httpbin.org/basic-auth/user/pass", auth=("user", "pass"), timeout=10)
print(r.status_code)  # 200

```

Basic auth is only base64, not encryption. Send it over HTTPS or not at all.

For bearer tokens, a small `AuthBase` subclass beats a hand-built header. It attaches cleanly to a session and keeps the token format in one place.

```python
import os
import requests
from requests.auth import AuthBase

class BearerAuth(AuthBase):
    def __init__(self, token):
        self.token = token

    def __call__(self, request):
        request.headers["Authorization"] = f"Bearer {self.token}"
        return request

r = requests.get("https://httpbin.org/bearer", auth=BearerAuth(os.environ["API_TOKEN"]), timeout=10)
print(r.json())  # {'authenticated': True, 'token': '...'}

```

There's a second reason to prefer `auth=`. If your `~/.netrc` has an entry for the host, requests applies it and overwrites an `Authorization` header you set by hand.

Keep tokens in environment variables, as above, and out of your repo. For OAuth 1 and 2 flows, the open-source `requests-oauthlib` package plugs into the same `auth=` argument.

## Set timeouts and handle errors

Requests has no default timeout. Without one, a server that accepts the connection and never answers holds your script indefinitely.

```python
import requests

requests.get("https://api.github.com", timeout=10)       # 10s for connect and 10s for read
requests.get("https://api.github.com", timeout=(3, 30))  # 3s to connect, 30s between bytes

```

The tuple form is worth the extra characters. A dead host should fail fast, while a slow report endpoint may need longer to send its first byte.

The read timeout is the gap allowed between bytes, not a limit on the whole download. A server that trickles one byte every nine seconds never trips `timeout=10`.

When you need a hard deadline, stream the body and check the clock yourself:

```python
import time
import requests

def get_with_deadline(url, deadline=30):
    start = time.monotonic()
    chunks = []
    with requests.get(url, stream=True, timeout=(5, 10)) as r:
        r.raise_for_status()
        for chunk in r.iter_content(64 * 1024):
            if time.monotonic() - start > deadline:
                raise TimeoutError(f"{url} took longer than {deadline}s")
            chunks.append(chunk)
    return b"".join(chunks)

```

The check only runs when a chunk arrives, so worst-case overrun is one read timeout. That's usually close enough.

### Catch the right exceptions

Every requests exception inherits from `RequestException`. The subclasses tell you whether retrying makes sense:

| Exception                 | Raised when                                   | Retry it?                             |
| ------------------------- | --------------------------------------------- | ------------------------------------- |
| ConnectTimeout            | TCP or TLS setup ran past the connect timeout | Yes. The server never saw the request |
| ReadTimeout               | The server accepted, then stalled             | Only for idempotent calls             |
| ConnectionError           | DNS failure, refused or reset connection      | Usually                               |
| HTTPError                 | raise\_for\_status() on a 4xx or 5xx          | 429 and 5xx yes, other 4xx no         |
| JSONDecodeError           | .json() on a non-JSON body                    | No. Read the body first               |
| MissingSchema, InvalidURL | The URL is malformed                          | No                                    |
| TooManyRedirects          | More than 30 redirects                        | No                                    |

`ConnectTimeout` inherits from both `ConnectionError` and `Timeout`. Put it first in your `except` chain, or the `ConnectionError` branch swallows it.

```python
import requests
from requests import exceptions as rex

def fetch_json(url):
    try:
        r = requests.get(url, timeout=(5, 15))
        r.raise_for_status()
        return r.json()
    except rex.ConnectTimeout:
        print("Couldn't connect in time")
    except rex.ReadTimeout:
        print("Server accepted the connection, then stalled")
    except rex.ConnectionError:
        print("DNS failure, refused, or reset")
    except rex.HTTPError as e:
        print(f"Bad status: {e.response.status_code}")
    except rex.JSONDecodeError:
        print(f"Expected JSON, got: {r.text[:200]!r}")
    except rex.RequestException as e:
        print(f"Request failed: {e}")

```

The JSON branch prints the first 200 characters of the body. When an API starts returning an HTML error page, that one line saves you twenty minutes.

## Reuse connections with a Session

Every bare `requests.get()` opens a fresh TCP connection and, for HTTPS, a fresh TLS handshake. A `Session` keeps finished connections in a pool and reuses them for the same host.

The saving is one or two network round trips per request. Against a server 100 ms away, that's 200 ms or more on every call after the first.

Sessions also carry headers, cookies, and auth across calls:

```python
import requests

repos = ["psf/requests", "encode/httpx", "aio-libs/aiohttp"]

with requests.Session() as session:
    session.headers.update({"Accept": "application/vnd.github+json", "User-Agent": "repo-audit/1.0"})
    for repo in repos:
        r = session.get(f"https://api.github.com/repos/{repo}", timeout=10)
        r.raise_for_status()
        print(repo, r.json()["stargazers_count"])

```

The `with` block closes pooled sockets on exit. In a long-running service, create one session at startup and reuse it for the life of the process.

One trap: `session.timeout = 5` looks right and does nothing, because Session has no timeout setting.

Python quietly creates a new attribute that requests never reads. When I pointed a session with that line at a 3-second endpoint, it waited the full 3 seconds.

The fix is an adapter, covered in the production session section.

## Retry failed requests automatically

Requests doesn't retry failed HTTP responses on its own. You attach a urllib3 `Retry` policy to an adapter, and the adapter handles the loop.

```python
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

retry = Retry(
    total=3,
    backoff_factor=1,                            # sleeps 0s, 2s, 4s between attempts
    status_forcelist=[429, 500, 502, 503, 504],
)
adapter = HTTPAdapter(max_retries=retry)

session = requests.Session()
session.mount("https://", adapter)
session.mount("http://", adapter)
r = session.get("https://httpbin.org/status/200", timeout=10)

```

The first retry fires immediately.

Later ones sleep `backoff_factor * 2 ** (n - 1)` seconds, capped at 120 by default. I timed three failed retries against an always-503 endpoint: 6.05 seconds total.

`Retry-After` is honored by default on 413, 429, and 503 responses. If a rate-limited API tells you to wait 30 seconds, urllib3 waits 30 seconds.

By default, only idempotent methods retry: GET, HEAD, PUT, DELETE, OPTIONS, and TRACE. POST is excluded on purpose.

Adding `"POST"` to `allowed_methods` can charge a card twice when the first attempt succeeded but the response got lost. Only do it if the API accepts idempotency keys.

When retries run out on a status code, you get `RetryError`, not the last response. Pass `raise_on_status=False` to `Retry` if you'd rather inspect the final 503 yourself.

Retries also change which exception reaches you. Once the policy gives up on read timeouts, requests raises `ConnectionError` with `Caused by ReadTimeoutError` in the message.

I hit this while testing: an `except ReadTimeout` branch never fired against a session with retries mounted. Catch `ConnectionError` there, or check `isinstance(e.args[0].reason, ReadTimeoutError)` if you need to tell them apart.

For deeper coverage of rate limits specifically, see [how to handle HTTP error 429](https://roundproxies.com/blog/http-error-429/).

## Use proxies with Python requests

Pass a dict that maps URL schemes to proxy URLs. Credentials go inside the URL.

```python
import os
import requests

proxy = os.environ["PROXY_URL"]  # e.g. http://user:pass@gate.example.com:8000
proxies = {"http": proxy, "https": proxy}

r = requests.get("https://httpbin.org/ip", proxies=proxies, timeout=15)
print(r.json())  # {'origin': '<the proxy exit IP>'}

```

The `"https"` key picks the proxy for HTTPS URLs. Its value still usually starts with `http://`, because requests opens a CONNECT tunnel through it.

Residential and datacenter endpoints from Roundproxies use the same `user:pass@host:port` format, so this dict works unchanged.

A few details catch people out:

- **SOCKS needs an extra.** Run `pip install "requests[socks]"`, then use `socks5h://` rather than `socks5://`. The `h` resolves DNS at the proxy, so lookups don't leak from your machine.
- **Environment variables count.** Requests reads `HTTP_PROXY`, `HTTPS_PROXY`, and `NO_PROXY`. On CI runners or corporate laptops, those can route traffic you didn't expect.
- **`session.proxies` can lose.** The requests docs warn that environment proxies override it. Pass `proxies=` on each call, or set `session.trust_env = False`.

Rotation, sticky sessions, and failover get the full treatment in the [Python requests proxy guide](https://roundproxies.com/blog/python-requests-proxy/).

## Stream large downloads

By default, requests loads the entire body into memory. A 4 GB file means 4 GB of RAM. Add `stream=True` and read in chunks:

```python
import requests

def download(url, path, chunk_size=256 * 1024):
    with requests.get(url, stream=True, timeout=(5, 30)) as r:
        r.raise_for_status()
        total = int(r.headers.get("Content-Length", 0))
        with open(path, "wb") as fh:
            for chunk in r.iter_content(chunk_size):
                fh.write(chunk)
                if total:
                    # raw.tell() counts bytes off the wire, same unit as Content-Length
                    print(f"\r{r.raw.tell() / total:.0%}", end="")
    print()

download("https://example.com/export.zip", "export.zip")

```

The progress math uses `r.raw.tell()` for a reason. `iter_content` hands you decompressed bytes, while `Content-Length` counts compressed bytes. Summing `len(chunk)` is the bug in most tutorials; details in the next section.

Chunked responses have no `Content-Length`, so the `if total` guard skips the percentage instead of dividing by zero.

For line-delimited feeds such as NDJSON exports, use `iter_lines()`:

```python
import json
import requests

with requests.get("https://example.com/events.ndjson", stream=True, timeout=(5, 60)) as r:
    for line in r.iter_lines():
        if line:  # skip keep-alive blank lines
            print(json.loads(line))

```

Always close streamed responses, which the `with` block does. An unread streamed response keeps its connection checked out of the pool.

## Make concurrent requests with threads

Requests blocks while it waits, but it releases the GIL during network I/O. A thread pool gets you real parallelism for fetching.

```python
from concurrent.futures import ThreadPoolExecutor
import requests
from requests.adapters import HTTPAdapter

WORKERS = 20
session = requests.Session()
# pool_maxsize must be >= WORKERS, or urllib3 throws finished sockets away
session.mount("https://", HTTPAdapter(pool_connections=4, pool_maxsize=WORKERS))

def fetch(url):
    r = session.get(url, timeout=10)
    return r.status_code, url

urls = [f"https://httpbin.org/anything/{i}" for i in range(100)]
with ThreadPoolExecutor(max_workers=WORKERS) as pool:
    for status, url in pool.map(fetch, urls):
        print(status, url)

```

The default pool holds 10 connections per host. When I ran 20 threads against a default session, urllib3 logged this 10 times:

```
Connection pool is full, discarding connection: 127.0.0.1. Connection pool size: 10

```

Nothing crashes, but every discarded socket means a new handshake later. Setting `pool_maxsize=20` brought the count to zero.

Sharing one session across threads works for stateless GETs, but requests doesn't promise it's thread-safe.

If workers log in or depend on their own cookies, give each thread its own session through `threading.local()`.

Past a few hundred concurrent requests, threads get expensive. That's the point to switch to httpx's async client or aiohttp.

## 5 requests snippets that quietly misbehave

These snippets show up across tutorials, Stack Overflow answers, and, until this rewrite, this page.

None of them raises an error. I ran each against a local test server with requests 2.33.1 and urllib3 2.6.3.

| Snippet                                          | What it claims       | What I measured                                                    | Fix                                   |
| ------------------------------------------------ | -------------------- | ------------------------------------------------------------------ | ------------------------------------- |
| Retry(total=3, backoff\_factor=1)                | Waits 1s, 2s, 4s     | Waits 0s, 2s, 4s. Three retries finished in 6.05s                  | Expect an immediate first retry       |
| Progress bar summing len(chunk)                  | Percent downloaded   | 86,957% on a gzip response: 230 bytes on the wire, 200,000 decoded | Divide r.raw.tell() by Content-Length |
| TimeoutHTTPAdapter that sets kwargs\["timeout"\] | A default timeout    | Replaced my explicit timeout=10 with its 1s default                | Only fill in when timeout is None     |
| session.timeout = 5                              | Session-wide timeout | Ignored. A 3s endpoint took the full 3.0s                          | Use an adapter default                |
| int(r.headers\["Retry-After"\])                  | Seconds to wait      | ValueError when the server sends an HTTP date                      | Parse both formats                    |

The last one crashes your error handler in the one moment you need it. `Retry-After` can legally be seconds or a date like `Wed, 21 Oct 2037 07:28:00 GMT`.

This helper handles both and falls back to a default:

```python
from datetime import datetime, timezone
from email.utils import parsedate_to_datetime

def retry_after_seconds(value, default=60.0):
    if value is None:
        return default
    value = value.strip()
    if value.isdigit():
        return float(value)
    try:
        when = parsedate_to_datetime(value)
    except (TypeError, ValueError):
        return default
    return max(0.0, (when - datetime.now(timezone.utc)).total_seconds())

```

Call it with `r.headers.get("Retry-After")`. A missing header returns the default instead of raising `KeyError`.

## Build a production-ready session

Everything above fits into two short pieces you can drop into any project. First, an adapter that applies a default timeout only when the caller didn't pass one:

```python
from requests.adapters import HTTPAdapter

class DefaultTimeoutAdapter(HTTPAdapter):
    """HTTPAdapter that fills in a timeout only when the caller omitted one."""

    def __init__(self, *args, timeout=(5, 30), **kwargs):
        self.timeout = timeout
        super().__init__(*args, **kwargs)

    def send(self, request, **kwargs):
        if kwargs.get("timeout") is None:
            kwargs["timeout"] = self.timeout
        return super().send(request, **kwargs)

```

Session always passes a `timeout` key to the adapter, set to `None` when you didn't specify one. Checking for `None` is what keeps explicit values intact.

Then a factory that wires retries, jitter, pool size, and an honest User-Agent into one session:

```python
import requests
from urllib3.util.retry import Retry

def make_session(user_agent, pool_size=10, retries=3):
    retry = Retry(
        total=retries,
        backoff_factor=0.5,
        backoff_jitter=0.3,  # spreads retries out when many workers fail together
        status_forcelist=[429, 500, 502, 503, 504],
    )
    adapter = DefaultTimeoutAdapter(max_retries=retry, pool_maxsize=pool_size)
    session = requests.Session()
    session.headers["User-Agent"] = user_agent
    session.mount("https://", adapter)
    session.mount("http://", adapter)
    return session

session = make_session("inventory-sync/2.0 (+ops@example.com)", pool_size=20)
r = session.get("https://httpbin.org/get")  # 5s connect, 30s read, 3 retries

```

`backoff_jitter` needs urllib3 2.0 or newer, which any current requests install already has. Without jitter, fifty workers that fail together retry together and hit the server in lockstep.

Build the session once, pass it around, and every call gets sane defaults without anyone remembering them.

## Troubleshooting common Python requests errors

These are the exact messages from my test runs, so you can match them against your traceback.

### `ModuleNotFoundError: No module named 'requests'`

**Why:** The package went into a different interpreter than the one running your script.

**Fix:** Print `sys.executable` inside the failing script, then run `<that path> -m pip install requests`.

### `MissingSchema: Invalid URL 'example.com': No scheme supplied. Perhaps you meant https://example.com?`

**Why:** Requests needs `http://` or `https://` in front of every URL.

**Fix:** Add the scheme. When URLs come from user input or a CSV, normalize them before the request.

### `ConnectionError: ... Max retries exceeded with url: / (Caused by NewConnectionError(... [Errno 111] Connection refused))`

**Why:** Nothing is listening on that host and port. "Max retries exceeded" is urllib3's wording and appears even when you configured zero retries.

**Fix:** Read the `Caused by` part. It tells you whether DNS failed, the port refused, or the connection reset. Check the host, port, and any proxy settings.

### `ReadTimeout: HTTPConnectionPool(...): Read timed out. (read timeout=1)`

**Why:** The connection opened, but the server went quiet for longer than your read timeout.

**Fix:** Raise the second value of the timeout tuple for slow endpoints. Retry only if the method is idempotent.

### `JSONDecodeError: Expecting value: line 1 column 1 (char 0)`

**Why:** The body isn't JSON. Usually it's an HTML error page, a CAPTCHA, or an empty 204.

**Fix:** Print `r.status_code` and `r.text[:200]` before calling `.json()`. The body tells you which of those it is.

### `SSLError: [SSL: CERTIFICATE_VERIFY_FAILED] certificate verify failed: unable to get local issuer certificate`

**Why:** The server's certificate chain doesn't lead to a CA in certifi's bundle. Corporate TLS inspection and self-signed internal services are the usual causes.

**Fix:** Run `python -m pip install --upgrade certifi`. For an internal CA, pass `verify="/path/to/company-ca.pem"`. Keep `verify=False` out of production.

### `RetryError: ... (Caused by ResponseError('too many 503 error responses'))`

**Why:** Your `Retry` policy ran out of attempts on a status in `status_forcelist`.

**Fix:** Catch `requests.exceptions.RetryError`, or set `raise_on_status=False` to get the final response back.

### `Connection pool is full, discarding connection`

**Why:** More threads than `pool_maxsize` are talking to one host.

**Fix:** Mount an `HTTPAdapter` with `pool_maxsize` at least equal to your worker count.

### 403 Forbidden on a page that loads in your browser

**Why:** The server flagged the request as automated. The `python-requests` User-Agent is the first suspect; the TLS fingerprint is the second.

**Fix:** Send browser headers as shown earlier. If a 403 survives that, the block is below the HTTP layer, and requests alone won't clear it.

## When requests is the wrong tool

Requests covers most HTTP work, but four situations call for something else:

- **The data appears only after JavaScript runs.** Requests fetches the HTML the server sends. Use browser automation, or find the JSON endpoint the page calls. The [Python web scraping guide](https://roundproxies.com/blog/web-scraping-python/) covers both routes.
- **You need hundreds of requests in flight.** Use httpx's `AsyncClient` or aiohttp.
- **The server wants HTTP/2.** Use httpx with the `http2` extra.
- **You're blocked on the TLS fingerprint.** The open-source curl\_cffi library impersonates browser handshakes behind a requests-style API.

For everything else, requests plus the session factory above is enough.

## FAQ

### Is requests part of the Python standard library?

No. Requests is a third-party package you install with `pip install requests`. The standard library's `urllib.request` handles HTTP without installs, but you encode parameters and parse responses by hand.

### What is the default timeout in Python requests?

There isn't one. Requests waits indefinitely for a response unless you pass `timeout=`.

Connection attempts eventually fail at the operating system level (roughly two minutes on Linux), but a stalled read can hang forever.

### How do I send JSON with Python requests?

Pass a dict to the `json=` argument: `requests.post(url, json={"key": "value"}, timeout=10)`. Requests serializes it and sets `Content-Type: application/json` for you.

### What's the difference between data and json in requests?

`data=` sends a dict as form-encoded fields, the format HTML forms use.

`json=` sends it as a JSON document. Pass only one of them; with both, requests sends the form data and silently drops the JSON.

### Is requests.Session thread-safe?

Not officially. Sharing one session for stateless GET requests works in practice if `pool_maxsize` matches your thread count. For logins or per-worker cookies, create one session per thread.

### Does Python requests support async?

No. Requests is synchronous and each call blocks its thread. For async code, httpx offers a nearly identical API through `httpx.AsyncClient`, and aiohttp suits very high concurrency.

## Wrapping up

The mental model is short: `requests.get()` for one-offs, a `Session` for anything repeated, and an adapter for timeouts, retries, and pool size.

Start by copying the `make_session()` factory into your project and routing every call through it.

It closes the timeout gap, retries the failures worth retrying, and sends a User-Agent you'd be comfortable explaining.

When your target starts blocking you, head to the [proxy guide for Python requests](https://roundproxies.com/blog/python-requests-proxy/). The [official requests documentation](https://requests.readthedocs.io/en/latest/) and the [urllib3 Retry reference](https://urllib3.readthedocs.io/en/stable/reference/urllib3.util.html) cover every argument this guide skipped.