> ## Content Index
> Fetch the complete content index at: https://roundproxies.com/blog/llms.txt
> Use this file to discover other available public pages before exploring further.

# Web Scraping with PHP in 2026: Step-by-step tutorial
- URL: https://roundproxies.com/blog/web-scraping-php/
- Published: 2025-10-29T17:52:13.000Z
- Updated: 2026-09-23T13:04:46.000Z
- Description: Learn web scraping with PHP in 2026. Build scrapers using cURL, Guzzle, and Panther with code examples. Handle JavaScript sites and anti-bot.
- Author: Marius Bernard
- Tags: Web Scraping, Web Scraping PHP, #dated-68fd1c4f358d710426c14657

PHP powers over 75% of websites with known server-side languages, so the scraper you build in this PHP web scraping tutorial can live inside the Laravel, Symfony, or WordPress codebase you already run. No second runtime to deploy, no glue layer to move rows from a Python script into your existing database.

Scraping is four steps in any language: fetch the page, parse the HTML, extract the fields you care about, store the result. PHP does all four with extensions that ship in the box, and its libraries take over at the point where hand-rolled code starts to hurt.

Two things changed recently that are worth knowing before you write a line of code. PHP 8.4 added a standards-compliant DOM API, and `Dom\HTMLDocument` parses HTML5 the way a browser does, which retires the error-suppression and encoding workarounds that older `DOMDocument` code is full of. And Goutte, for years the default answer for crawling in PHP, is abandoned. Its functionality now lives in Symfony's `HttpBrowser`, which is the path to take for anything involving links, forms, or logged-in sessions.

Each section builds on the one before it: a raw cURL request to see the mechanics, Guzzle and Symfony DomCrawler for work you intend to keep, Symfony Panther for pages that only exist after JavaScript runs, and the reliability details (timeouts, retries, rate limiting) that decide whether a scraper still works next month.

## What is Web Scraping with PHP?

Web scraping with PHP is the process of programmatically extracting data from websites using PHP scripts. You send HTTP requests to target URLs, receive HTML responses, and parse that HTML to pull out the specific data you need.

PHP handles this well because it was built for the web. It has native cURL support, excellent string manipulation functions, and libraries specifically designed for DOM traversal and CSS selector matching.

## Why PHP for web scraping in 2026

PHP gets overlooked in scraping discussions because Python owns the conversation with BeautifulSoup and Scrapy. The case for PHP is narrower, and it holds up: when the scraped data ends up inside a PHP application anyway, scraping it in PHP removes a layer of glue. No second runtime on the server, no Python virtualenv living next to your Laravel or WordPress install, no cron job shelling out to a language the rest of the team does not maintain. You fetch a page, parse it, and write the rows straight into the same MySQL database your app already reads from.

### What changed in PHP 8.3 and 8.4

PHP 8.3 and 8.4 brought performance improvements once you are parsing thousands of documents. JIT compilation speeds up the tight loops in parsing and string handling compared with PHP 7.x.

The bigger change for scrapers is parsing. PHP 8.4 ships a standards-compliant DOM API, and `Dom\HTMLDocument` parses HTML5 the way a browser does. That removes the two rituals every older PHP scraper is built on: suppressing the warning storm that `DOMDocument::loadHTML()` throws on real-world markup, and hand-patching encoding declarations so accented characters survive. Older `DOMDocument` code keeps working, and the examples further down use it because plenty of production servers are still on earlier versions. If you control the runtime and it is on 8.4 or newer, reach for the new API first.

### The 2026 library stack

Four tools cover almost everything, and picking the lightest one that does the job is most of the skill:

- **cURL** ships with nearly every PHP installation. Verbose, no dependencies, full control over headers, timeouts, and redirects. Fine for a single-file script.
- **Guzzle** is the HTTP client for anything larger. Clean handling of POST bodies, JSON, cookies, and redirects, plus concurrent requests and retry middleware you would otherwise write yourself.
- **Symfony DomCrawler** queries the parsed document with XPath, or with CSS selectors once you also install `symfony/css-selector`.
- **Symfony HttpBrowser** handles stateful workflows: follow a link, submit a login form, keep the session cookie across requests. Most PHP web scraping tutorials still point at Goutte for this, but Goutte is deprecated and its functionality was folded into HttpBrowser, so new projects should start there.
- **Symfony Panther** drives a real browser and belongs at the bottom of the list. Use it only after you have confirmed the data genuinely does not exist in the raw HTML, because it costs you memory, startup time, and a browser binary on the server.

### Where PHP fits, and where it runs out of road

PHP is a comfortable choice for scheduled jobs on a cron timer, WordPress plugins and Laravel console commands that enrich existing records, and small to medium crawls that finish in minutes. Composer makes dependency management painless, memory handling improved substantially in recent versions, and deployment to cheap shared hosting is realistic in a way it is not for a Python worker fleet.

The gap shows up at scale. PHP has no Scrapy equivalent, so the crawler architecture is yours to build: a frontier of URLs waiting to be visited, a visited set so you do not fetch the same page twice, persistence for that queue (a small SQLite or MySQL table is enough) so an interrupted run resumes instead of restarting, plus concurrency and exponential backoff on failures. Long-running scripts also collide with shared hosting limits, where `max_execution_time` will kill a crawl mid-run unless you shard the work into batches.

For massive distributed scraping across rotating proxies, expect to write infrastructure that Python frameworks hand you for free. For everything below that line, PHP does the job and keeps your stack to one language.

## Setting Up Your PHP Environment

Before writing any scraping code, confirm your PHP installation meets the requirements.

Open your terminal and check the PHP version:

```bash
php --version

```

You need PHP 8.1 or higher. PHP 8.3+ is recommended for best performance.

Next, verify Composer is installed:

```bash
composer --version

```

If Composer isn't available, install it from [getcomposer.org](https://getcomposer.org/download/).

Create a new project directory and initialize Composer:

```bash
mkdir php-scraper
cd php-scraper
composer init --no-interaction --require="php >=8.1"

```

This creates your `composer.json` file. You're ready to add libraries as needed.

## Method 1: Native cURL Approach

cURL comes bundled with PHP. It's the foundation of most HTTP operations in the language.

Start with a basic request. This fetches the raw HTML from a target URL:

```php
<?php
// Initialize a cURL session
$ch = curl_init();

// Set the target URL
curl_setopt($ch, CURLOPT_URL, 'https://books.toscrape.com');

// Return the response instead of printing it
curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);

// Follow redirects automatically
curl_setopt($ch, CURLOPT_FOLLOWLOCATION, true);

// Set a timeout to avoid hanging
curl_setopt($ch, CURLOPT_TIMEOUT, 30);

// Execute and store the response
$html = curl_exec($ch);

// Check for errors
if (curl_errno($ch)) {
    echo 'cURL Error: ' . curl_error($ch);
} else {
    echo "Fetched " . strlen($html) . " bytes\n";
}

// Close the session
curl_close($ch);

```

This returns raw HTML as a string. To extract specific data, combine cURL with DOMDocument.

### Parsing HTML with DOMDocument

DOMDocument is PHP's built-in XML/HTML parser. It converts HTML strings into traversable document objects.

```php
<?php
$ch = curl_init();
curl_setopt($ch, CURLOPT_URL, 'https://books.toscrape.com');
curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
curl_setopt($ch, CURLOPT_FOLLOWLOCATION, true);
$html = curl_exec($ch);
curl_close($ch);

// Suppress warnings from malformed HTML
libxml_use_internal_errors(true);

// Create DOM document and load HTML
$dom = new DOMDocument();
$dom->loadHTML($html);

// Use XPath for element selection
$xpath = new DOMXPath($dom);

// Find all book titles
$titles = $xpath->query('//article[@class="product_pod"]//h3/a/@title');

foreach ($titles as $title) {
    echo $title->nodeValue . "\n";
}

```

The XPath query locates `<article>` elements with the `product_pod` class, then drills down to the `<h3>` link's title attribute.

DOMDocument handles most real-world HTML despite validator errors. The `libxml_use_internal_errors(true)` line prevents warning spam.

### Adding Request Headers

Websites check headers to identify automated requests. Set a realistic User-Agent and other headers:

```php
<?php
$ch = curl_init();

curl_setopt_array($ch, [
    CURLOPT_URL => 'https://books.toscrape.com',
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_FOLLOWLOCATION => true,
    CURLOPT_TIMEOUT => 30,
    CURLOPT_HTTPHEADER => [
        'User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36',
        'Accept: text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8',
        'Accept-Language: en-US,en;q=0.5',
        'Accept-Encoding: gzip, deflate',
        'Connection: keep-alive',
    ],
    CURLOPT_ENCODING => 'gzip',
]);

$html = curl_exec($ch);
$httpCode = curl_getinfo($ch, CURLINFO_HTTP_CODE);

echo "Status: $httpCode\n";
echo "Size: " . strlen($html) . " bytes\n";

curl_close($ch);

```

The `CURLOPT_ENCODING` option tells cURL to decompress gzip responses automatically.

## Method 2: Using Guzzle HTTP Client

Guzzle wraps cURL in a modern, object-oriented interface. It handles cookies, sessions, and concurrent requests more cleanly than raw cURL.

Install Guzzle via Composer:

```bash
composer require guzzlehttp/guzzle

```

Basic usage looks like this:

```php
<?php
require 'vendor/autoload.php';

use GuzzleHttp\Client;
use GuzzleHttp\Exception\RequestException;

$client = new Client([
    'timeout' => 30,
    'verify' => false,
    'headers' => [
        'User-Agent' => 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) Chrome/120.0.0.0',
    ],
]);

try {
    $response = $client->get('https://books.toscrape.com');
    $html = $response->getBody()->getContents();
    
    echo "Status: " . $response->getStatusCode() . "\n";
    echo "Fetched " . strlen($html) . " bytes\n";
    
} catch (RequestException $e) {
    echo "Request failed: " . $e->getMessage() . "\n";
}

```

Guzzle's exception handling makes error management cleaner than checking `curl_errno()`.

### Concurrent Requests with Guzzle

Scraping multiple pages sequentially wastes time. Guzzle supports concurrent requests through its Pool class:

```php
<?php
require 'vendor/autoload.php';

use GuzzleHttp\Client;
use GuzzleHttp\Pool;
use GuzzleHttp\Psr7\Request;

$client = new Client(['timeout' => 30]);

$urls = [
    'https://books.toscrape.com/catalogue/page-1.html',
    'https://books.toscrape.com/catalogue/page-2.html',
    'https://books.toscrape.com/catalogue/page-3.html',
    'https://books.toscrape.com/catalogue/page-4.html',
    'https://books.toscrape.com/catalogue/page-5.html',
];

$requests = function () use ($urls) {
    foreach ($urls as $url) {
        yield new Request('GET', $url);
    }
};

$results = [];

$pool = new Pool($client, $requests(), [
    'concurrency' => 5,
    'fulfilled' => function ($response, $index) use (&$results, $urls) {
        $results[$urls[$index]] = $response->getBody()->getContents();
        echo "Completed: " . $urls[$index] . "\n";
    },
    'rejected' => function ($reason, $index) use ($urls) {
        echo "Failed: " . $urls[$index] . " - " . $reason . "\n";
    },
]);

$promise = $pool->promise();
$promise->wait();

echo "Scraped " . count($results) . " pages\n";

```

The `concurrency` parameter controls how many requests run simultaneously. Setting it too high can trigger rate limits or IP blocks.

## Method 3: Symfony DomCrawler for Parsing

DomCrawler provides jQuery-like syntax for HTML parsing. It's more intuitive than raw XPath for most developers.

Install it with Composer:

```bash
composer require symfony/dom-crawler symfony/css-selector

```

The CSS Selector component enables familiar selectors like `.class` and `#id`.

```php
<?php
require 'vendor/autoload.php';

use GuzzleHttp\Client;
use Symfony\Component\DomCrawler\Crawler;

$client = new Client();
$response = $client->get('https://books.toscrape.com');
$html = $response->getBody()->getContents();

$crawler = new Crawler($html);

// Extract all book data using CSS selectors
$books = $crawler->filter('article.product_pod')->each(function (Crawler $node) {
    return [
        'title' => $node->filter('h3 a')->attr('title'),
        'price' => $node->filter('.price_color')->text(),
        'stock' => $node->filter('.availability')->text(),
        'rating' => $node->filter('.star-rating')->attr('class'),
    ];
});

// Display results
foreach ($books as $book) {
    echo $book['title'] . ' - ' . $book['price'] . "\n";
}

```

The `filter()` method accepts CSS selectors. The `each()` method iterates over matched elements, returning an array of extracted data.

### Navigating Between Pages

DomCrawler can extract links for pagination:

```php
<?php
require 'vendor/autoload.php';

use GuzzleHttp\Client;
use Symfony\Component\DomCrawler\Crawler;

$client = new Client();
$baseUrl = 'https://books.toscrape.com/catalogue/';
$currentUrl = $baseUrl . 'page-1.html';
$allBooks = [];

while ($currentUrl) {
    echo "Scraping: $currentUrl\n";
    
    $response = $client->get($currentUrl);
    $crawler = new Crawler($response->getBody()->getContents());
    
    // Extract books from current page
    $crawler->filter('article.product_pod')->each(function ($node) use (&$allBooks) {
        $allBooks[] = [
            'title' => $node->filter('h3 a')->attr('title'),
            'price' => $node->filter('.price_color')->text(),
        ];
    });
    
    // Find next page link
    try {
        $nextLink = $crawler->filter('.next a')->attr('href');
        $currentUrl = $baseUrl . $nextLink;
    } catch (\Exception $e) {
        $currentUrl = null; // No more pages
    }
    
    // Be respectful with delays
    usleep(500000); // 0.5 second delay
}

echo "Total books scraped: " . count($allBooks) . "\n";

```

The `usleep()` call adds a half-second delay between requests. This prevents hammering the server and reduces your chance of getting blocked.

## Method 4: Headless Browsers with Panther

Static HTML scrapers fail on JavaScript-heavy websites. Modern sites render content dynamically, meaning the initial HTML contains no useful data.

Symfony Panther controls real browsers (Chrome or Firefox) programmatically. It waits for JavaScript to execute and renders the complete page.

Install Panther:

```bash
composer require symfony/panther

```

You also need ChromeDriver or GeckoDriver. On macOS:

```bash
brew install chromedriver

```

On Windows with Chocolatey:

```bash
choco install chromedriver

```

Basic Panther usage:

```php
<?php
require 'vendor/autoload.php';

use Symfony\Component\Panther\Client;

// Create a Chrome client
$client = Client::createChromeClient();

// Navigate to a JavaScript-heavy site
$crawler = $client->request('GET', 'https://quotes.toscrape.com/js/');

// Wait for content to load
$client->waitFor('.quote');

// Extract quotes after JavaScript execution
$quotes = $crawler->filter('.quote')->each(function ($node) {
    return [
        'text' => $node->filter('.text')->text(),
        'author' => $node->filter('.author')->text(),
    ];
});

foreach ($quotes as $quote) {
    echo $quote['author'] . ': ' . $quote['text'] . "\n\n";
}

// Always close the browser
$client->quit();

```

The `waitFor()` method pauses execution until the specified CSS selector appears in the DOM. This handles async content loading.

### Interacting with Page Elements

Panther can click buttons, fill forms, and trigger JavaScript events:

```php
<?php
require 'vendor/autoload.php';

use Symfony\Component\Panther\Client;

$client = Client::createChromeClient();
$crawler = $client->request('GET', 'https://quotes.toscrape.com/js/');

$allQuotes = [];

while (true) {
    $client->waitFor('.quote');
    
    // Extract quotes from current page state
    $crawler->filter('.quote')->each(function ($node) use (&$allQuotes) {
        $allQuotes[] = [
            'text' => $node->filter('.text')->text(),
            'author' => $node->filter('.author')->text(),
        ];
    });
    
    // Try clicking the Next button
    try {
        $client->clickLink('Next');
        usleep(1000000); // Wait 1 second for page load
        $crawler = $client->refreshCrawler();
    } catch (\Exception $e) {
        break; // No more pages
    }
}

echo "Scraped " . count($allQuotes) . " quotes\n";

$client->quit();

```

Panther is resource-intensive. Each instance runs a full browser process. Use it only when static methods won't work.

## Handling Anti-Bot Detection

Websites deploy various techniques to [block scrapers](https://roundproxies.com/blog/why-your-cloudflare-bypass-always-fails/). Here's how to work around common obstacles.

### Rotating User Agents

Cycling through different User-Agent strings makes requests appear more natural:

```php
<?php
$userAgents = [
    'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 Chrome/120.0.0.0 Safari/537.36',
    'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 Chrome/120.0.0.0 Safari/537.36',
    'Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:121.0) Gecko/20100101 Firefox/121.0',
    'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 Safari/17.2',
    'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 Chrome/120.0.0.0 Safari/537.36',
];

function getRandomUserAgent() {
    global $userAgents;
    return $userAgents[array_rand($userAgents)];
}

```

### Using Proxies

Proxies route requests through different IP addresses. This prevents IP-based blocking when scraping at scale.

```php
<?php
$ch = curl_init();

curl_setopt_array($ch, [
    CURLOPT_URL => 'https://httpbin.org/ip',
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_PROXY => 'http://proxy.example.com:8080',
    CURLOPT_PROXYUSERPWD => 'username:password',
]);

$response = curl_exec($ch);
echo $response; // Shows the proxy IP

curl_close($ch);

```

For rotating proxies, services like Roundproxies.com offer residential and datacenter proxy pools that automatically cycle IPs between requests. This lowers detection rates compared to single-IP scraping.

With Guzzle, proxy configuration looks like this:

```php
<?php
$client = new GuzzleHttp\Client([
    'proxy' => 'http://username:password@proxy.example.com:8080',
    'timeout' => 30,
]);

```

### Implementing Request Delays

Hammering a server with rapid requests triggers rate limiting. Space out your requests:

```php
<?php
class RateLimiter
{
    private array $lastRequest = [];
    
    public function wait(string $domain, float $minDelay = 1.0): void
    {
        if (isset($this->lastRequest[$domain])) {
            $elapsed = microtime(true) - $this->lastRequest[$domain];
            if ($elapsed < $minDelay) {
                $sleepTime = ($minDelay - $elapsed) * 1000000;
                usleep((int)$sleepTime);
            }
        }
        $this->lastRequest[$domain] = microtime(true);
    }
}

// Usage
$limiter = new RateLimiter();

foreach ($urls as $url) {
    $domain = parse_url($url, PHP_URL_HOST);
    $limiter->wait($domain, 1.5); // 1.5 seconds minimum between requests
    
    // Make request...
}

```

### Handling Cookies and Sessions

Some sites require maintaining session state:

```php
<?php
$client = new GuzzleHttp\Client([
    'cookies' => true, // Enable cookie jar
]);

// First request establishes session
$client->get('https://example.com/');

// Subsequent requests maintain cookies
$response = $client->get('https://example.com/protected-page');

```

Guzzle's cookie jar automatically stores and sends cookies between requests.

## Saving Scraped Data

Extracted data needs storage. Common formats include CSV, JSON, and databases.

### Writing to CSV

PHP's `fputcsv()` handles CSV creation:

```php
<?php
$books = [
    ['title' => 'Book One', 'price' => '£51.77'],
    ['title' => 'Book Two', 'price' => '£53.74'],
];

$file = fopen('books.csv', 'w');

// Write header row
fputcsv($file, ['Title', 'Price']);

// Write data rows
foreach ($books as $book) {
    fputcsv($file, [$book['title'], $book['price']]);
}

fclose($file);

echo "Saved " . count($books) . " books to CSV\n";

```

### Writing to JSON

JSON output preserves nested structures better:

```php
<?php
$books = [
    ['title' => 'Book One', 'price' => '£51.77', 'details' => ['pages' => 320]],
    ['title' => 'Book Two', 'price' => '£53.74', 'details' => ['pages' => 256]],
];

$json = json_encode($books, JSON_PRETTY_PRINT | JSON_UNESCAPED_UNICODE);

file_put_contents('books.json', $json);

echo "Saved JSON data\n";

```

### Database Storage with PDO

For larger scraping operations, database storage makes querying easier:

```php
<?php
$pdo = new PDO('sqlite:scraper.db');

$pdo->exec('CREATE TABLE IF NOT EXISTS books (
    id INTEGER PRIMARY KEY,
    title TEXT NOT NULL,
    price TEXT,
    scraped_at DATETIME DEFAULT CURRENT_TIMESTAMP
)');

$stmt = $pdo->prepare('INSERT INTO books (title, price) VALUES (?, ?)');

foreach ($books as $book) {
    $stmt->execute([$book['title'], $book['price']]);
}

echo "Inserted " . count($books) . " records\n";

```

SQLite works well for local scraping scripts. For production, use MySQL or PostgreSQL.

## Common Errors and Fixes

### SSL Certificate Errors

Some servers have misconfigured SSL. Bypass verification (only for trusted targets):

```php
curl_setopt($ch, CURLOPT_SSL_VERIFYPEER, false);
curl_setopt($ch, CURLOPT_SSL_VERIFYHOST, 0);

```

With Guzzle:

```php
$client = new Client(['verify' => false]);

```

### Memory Issues with Large Scrapes

Processing thousands of pages can exhaust memory. Force garbage collection periodically:

```php
<?php
$counter = 0;

foreach ($urls as $url) {
    // Scrape and process...
    
    $counter++;
    if ($counter % 100 === 0) {
        gc_collect_cycles();
        echo "Memory: " . (memory_get_usage(true) / 1024 / 1024) . " MB\n";
    }
}

```

Also unset large variables when done with them:

```php
$html = $client->get($url)->getBody()->getContents();
$data = parseHtml($html);
unset($html); // Free memory immediately

```

### Timeout Errors

Slow servers need longer timeouts:

```php
curl_setopt($ch, CURLOPT_TIMEOUT, 60); // 60 second timeout
curl_setopt($ch, CURLOPT_CONNECTTIMEOUT, 10); // 10 second connection timeout

```

### Character Encoding Issues

Force UTF-8 encoding when parsing:

```php
$html = mb_convert_encoding($html, 'HTML-ENTITIES', 'UTF-8');
$dom->loadHTML($html);

```

## Complete Web Scraper Class

This production-ready scraper class combines everything covered above:

```php
<?php
require 'vendor/autoload.php';

use GuzzleHttp\Client;
use GuzzleHttp\Exception\RequestException;
use Symfony\Component\DomCrawler\Crawler;

class WebScraper
{
    private Client $client;
    private array $userAgents;
    private array $lastRequest = [];
    private float $minDelay;
    
    public function __construct(float $minDelay = 1.0, ?string $proxy = null)
    {
        $this->minDelay = $minDelay;
        
        $this->userAgents = [
            'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 Chrome/120.0.0.0 Safari/537.36',
            'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 Chrome/120.0.0.0 Safari/537.36',
            'Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:121.0) Gecko/20100101 Firefox/121.0',
        ];
        
        $config = [
            'timeout' => 30,
            'verify' => false,
            'cookies' => true,
        ];
        
        if ($proxy) {
            $config['proxy'] = $proxy;
        }
        
        $this->client = new Client($config);
    }
    
    private function getRandomUserAgent(): string
    {
        return $this->userAgents[array_rand($this->userAgents)];
    }
    
    private function respectRateLimit(string $domain): void
    {
        if (isset($this->lastRequest[$domain])) {
            $elapsed = microtime(true) - $this->lastRequest[$domain];
            if ($elapsed < $this->minDelay) {
                usleep((int)(($this->minDelay - $elapsed) * 1000000));
            }
        }
        $this->lastRequest[$domain] = microtime(true);
    }
    
    public function fetch(string $url, int $retries = 3): ?string
    {
        $domain = parse_url($url, PHP_URL_HOST);
        $this->respectRateLimit($domain);
        
        $attempt = 0;
        while ($attempt < $retries) {
            try {
                $response = $this->client->get($url, [
                    'headers' => [
                        'User-Agent' => $this->getRandomUserAgent(),
                        'Accept' => 'text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8',
                        'Accept-Language' => 'en-US,en;q=0.5',
                    ],
                ]);
                
                return $response->getBody()->getContents();
                
            } catch (RequestException $e) {
                $attempt++;
                echo "Attempt $attempt failed: " . $e->getMessage() . "\n";
                
                if ($attempt < $retries) {
                    sleep(pow(2, $attempt)); // Exponential backoff
                }
            }
        }
        
        return null;
    }
    
    public function parse(string $html, string $selector): array
    {
        $crawler = new Crawler($html);
        
        return $crawler->filter($selector)->each(function (Crawler $node) {
            return $node->text();
        });
    }
    
    public function scrapeWithCallback(
        string $url, 
        string $itemSelector, 
        callable $extractor
    ): array {
        $html = $this->fetch($url);
        
        if (!$html) {
            return [];
        }
        
        $crawler = new Crawler($html);
        
        return $crawler->filter($itemSelector)->each($extractor);
    }
}

// Usage example
$scraper = new WebScraper(minDelay: 1.5);

$books = $scraper->scrapeWithCallback(
    'https://books.toscrape.com',
    'article.product_pod',
    function (Crawler $node) {
        return [
            'title' => $node->filter('h3 a')->attr('title'),
            'price' => $node->filter('.price_color')->text(),
        ];
    }
);

print_r($books);

```

This class handles rate limiting, user agent rotation, retries with exponential backoff, and clean callback-based extraction. Extend it for your specific needs.

## Scraping APIs vs HTML Parsing

Not all data requires HTML parsing. Many websites expose APIs that return clean JSON. Check your browser's Network tab while browsing the target site.

If you find API endpoints, they're usually easier to work with:

```php
<?php
$client = new GuzzleHttp\Client();

$response = $client->get('https://api.example.com/products', [
    'query' => [
        'page' => 1,
        'limit' => 50,
    ],
    'headers' => [
        'Accept' => 'application/json',
    ],
]);

$data = json_decode($response->getBody()->getContents(), true);

foreach ($data['products'] as $product) {
    echo $product['name'] . ' - $' . $product['price'] . "\n";
}

```

JSON responses skip the parsing complexity entirely. The data arrives structured and ready to use.

Look for GraphQL endpoints too. They often expose more data than visible on the page:

```php
<?php
$client = new GuzzleHttp\Client();

$query = <<<GRAPHQL
{
    products(first: 50) {
        edges {
            node {
                title
                price
                description
            }
        }
    }
}
GRAPHQL;

$response = $client->post('https://example.com/graphql', [
    'json' => ['query' => $query],
]);

$result = json_decode($response->getBody()->getContents(), true);

```

GraphQL lets you request exactly the fields you need, reducing response size and processing time.

## Performance Optimization Tips

Large scraping jobs need optimization to finish in reasonable time.

### Connection Reuse

Create one cURL handle and reuse it for multiple requests:

```php
<?php
$ch = curl_init();

curl_setopt_array($ch, [
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_FOLLOWLOCATION => true,
    CURLOPT_TIMEOUT => 30,
]);

$urls = ['https://example.com/page1', 'https://example.com/page2'];
$results = [];

foreach ($urls as $url) {
    curl_setopt($ch, CURLOPT_URL, $url);
    $results[$url] = curl_exec($ch);
    usleep(500000);
}

curl_close($ch);

```

Reusing handles avoids TCP handshake overhead for each request.

### Async Processing with ReactPHP

For high-throughput scraping, ReactPHP provides true async capabilities:

```bash
composer require react/http react/event-loop

```

```php
<?php
require 'vendor/autoload.php';

use React\EventLoop\Loop;
use React\Http\Browser;

$browser = new Browser();
$urls = [
    'https://books.toscrape.com/catalogue/page-1.html',
    'https://books.toscrape.com/catalogue/page-2.html',
    'https://books.toscrape.com/catalogue/page-3.html',
];

$promises = [];

foreach ($urls as $url) {
    $promises[$url] = $browser->get($url)->then(
        function ($response) use ($url) {
            echo "Completed: $url\n";
            return (string) $response->getBody();
        },
        function ($error) use ($url) {
            echo "Failed: $url - " . $error->getMessage() . "\n";
            return null;
        }
    );
}

React\Promise\all($promises)->then(function ($results) {
    echo "All done. Got " . count(array_filter($results)) . " pages\n";
});

Loop::run();

```

ReactPHP runs multiple requests truly in parallel within a single PHP process.

### Caching Responses

Avoid re-fetching unchanged pages:

```php
<?php
class CachedScraper
{
    private string $cacheDir;
    private int $cacheTime;
    
    public function __construct(string $cacheDir = './cache', int $cacheTime = 3600)
    {
        $this->cacheDir = $cacheDir;
        $this->cacheTime = $cacheTime;
        
        if (!is_dir($cacheDir)) {
            mkdir($cacheDir, 0755, true);
        }
    }
    
    public function fetch(string $url): string
    {
        $cacheFile = $this->cacheDir . '/' . md5($url) . '.html';
        
        if (file_exists($cacheFile) && (time() - filemtime($cacheFile)) < $this->cacheTime) {
            return file_get_contents($cacheFile);
        }
        
        $html = file_get_contents($url);
        file_put_contents($cacheFile, $html);
        
        return $html;
    }
}

```

Caching is essential during development when you're testing selectors repeatedly.

## Scheduled Scraping with Cron

Production scrapers usually run on schedules. Set up a cron job to run your PHP script:

```bash
# Edit crontab
crontab -e

# Run scraper every 6 hours
0 */6 * * * /usr/bin/php /path/to/scraper.php >> /var/log/scraper.log 2>&1

```

For more complex scheduling, Laravel's task scheduler or Symfony Console Commands provide better management.

Basic standalone scheduler pattern:

```php
<?php
// scraper.php
$startTime = date('Y-m-d H:i:s');
echo "[$startTime] Starting scrape job\n";

try {
    // Your scraping logic here
    $count = runScraper();
    
    $endTime = date('Y-m-d H:i:s');
    echo "[$endTime] Completed. Scraped $count items.\n";
    
} catch (Exception $e) {
    $endTime = date('Y-m-d H:i:s');
    echo "[$endTime] Error: " . $e->getMessage() . "\n";
    exit(1);
}

```

Log output and errors for debugging. Consider using Monolog for structured logging in larger projects.

## FAQ

### Can PHP handle JavaScript-rendered websites?

Yes, through headless browser libraries like Symfony Panther. Panther controls real Chrome or Firefox instances, waits for JavaScript to execute, then exposes the rendered DOM for parsing. It's resource-intensive compared to static methods but necessary for modern SPAs.

### How do I avoid getting blocked while scraping?

Use realistic headers including proper User-Agent strings. Add delays between requests (1-2 seconds minimum). Consider rotating proxies through services like Roundproxies.com to distribute requests across IPs. Don't run hundreds of parallel requests against a single domain.

### Is web scraping with PHP legal?

The legality depends on what you're scraping and how you use the data. Publicly available data is generally fair game. However, violating Terms of Service, accessing private data, or overloading servers can create legal issues. Always check the target site's robots.txt and terms of use.

### Which PHP version is best for scraping?

PHP 8.3 or 8.4 offers the best performance thanks to JIT compilation. The minimum recommended version is 8.1 since older versions lack security updates and modern language features.

### How does PHP compare to Python for web scraping?

Python has a larger scraping ecosystem with tools like Scrapy and BeautifulSoup. It handles concurrency better at scale. PHP is simpler when you already run PHP infrastructure and need moderate scraping capabilities. For enterprise-level distributed scraping, Python typically wins. For WordPress plugins or Laravel integrations, PHP makes more sense.

## Final Thoughts

Web scraping with PHP works well when your project already uses PHP or when you need simple, reliable data extraction. Libraries like Guzzle, DomCrawler, and Panther handle most common scenarios.

For static websites, the Guzzle + DomCrawler combination offers the best balance of simplicity and power. Add Panther when JavaScript rendering becomes necessary.

Keep these principles in mind:

- Respect robots.txt and rate limits
- Use delays between requests to avoid overwhelming servers
- Rotate headers and consider proxies for larger operations
- Store data incrementally to avoid losing progress on failures
- Handle errors with retry logic

PHP may not be the trendiest choice for web scraping in 2026, but it remains practical. If you're comfortable with the language and your requirements are moderate, there's no reason to add Python complexity to your stack.

Start with cURL and DOMDocument for simple tasks. Graduate to Guzzle and Symfony components as needs grow. Save headless browsers for sites that truly require them.