Puppeteer

How to get elements in Puppeteer: class, ID, text, XPath

Most broken Puppeteer scripts die on the line that fetches an element. page.$() returns null, the next .click() throws, and the stack trace blames the wrong line.

Older tutorials make it worse. page.$x() was removed in Puppeteer 22, and Puppeteer 25 ships as ESM only, so a lot of the snippets you'll find online no longer run.

This guide shows how to get an element in Puppeteer 25 by class, ID, attribute, text, XPath, and ARIA role, plus inside shadow DOM and iframes, all with current code.

How do you get an element in Puppeteer?

To get an element in Puppeteer, call page.$(selector) for the first match or page.$$(selector) for every match. Both return ElementHandles you can click or type into. Use page.$eval() or page.$$eval() when you only need text or attributes. On JavaScript-rendered pages, use page.locator() or page.waitForSelector() first, because the $ methods never wait.

Every one of these accepts a CSS selector by default.

Puppeteer also layers its own syntax on top, so the same methods can match by visible text, XPath, or accessibility role without any page.evaluate() code.

The practical choice comes down to one question: do you need a handle to act on, or data to save?

Handles are live references into Chrome, while data is a plain copy that can't go stale.

Which Puppeteer method to use: quick reference

Bookmark this table. The "If nothing matches" column explains most of the confusing bugs people hit, because each method fails in a different way.

Method Returns If nothing matches Waits? Use it for
page.$(sel) ElementHandle Returns null No One element you'll click, type into, or screenshot
page.$$(sel) ElementHandle[] Returns [] No A few elements you'll interact with one by one
page.$eval(sel, fn) Whatever fn returns Throws an error No Text, href, or value from one element
page.$$eval(sel, fn) Whatever fn returns Runs fn on an empty array No Bulk extraction from lists and tables
page.waitForSelector(sel) ElementHandle Throws TimeoutError (30s default) Yes Late-rendering content you need as a handle
page.locator(sel) Locator Throws TimeoutError Yes, plus visibility and stability checks Clicking, filling, hovering

My default split: locators for anything interactive, $$eval for extraction, and $/$$ only when I need a handle for something locators don't cover.

Prerequisites

You need three things before any of the code below will run:

  • Node.js 22 or newer. Puppeteer 25 raised the minimum version.
  • Puppeteer itself: npm i puppeteer, which also downloads a matching Chrome for Testing build.
  • ES modules: name your files .mjs or add "type": "module" to package.json. The Puppeteer changelog lists the move to ESM-only packages under 25.0.0.

Here's the skeleton every example plugs into. The examples target books.toscrape.com and quotes.toscrape.com, two public sandboxes built for scraping practice.

// setup.mjs
import puppeteer from 'puppeteer';

const browser = await puppeteer.launch();
const page = await browser.newPage();
await page.goto('https://books.toscrape.com/');

// ...queries go here...

await browser.close();

Top-level await works in ES modules, so the old (async () => { ... })() wrapper is gone.

If Node complains about loading an ES module through require(), switch the file to import syntax.

Step 1: Get the first match with page.$()

page.$() runs document.querySelector() inside Chrome and hands back a handle to the first match. On books.toscrape.com, each book sits in an article.product_pod card.

const firstBook = await page.$('article.product_pod');

if (!firstBook) {
  throw new Error('No product cards on the page');
}

// Handles run their own queries, scoped to their subtree
const title = await firstBook.$eval('h3 a', a => a.getAttribute('title'));
console.log(title);

await firstBook.dispose(); // release the reference inside Chrome

Two details matter here. The null check exists because page.$() returns null on a miss instead of throwing, so an unchecked handle crashes on the next line.

The title comes from the title attribute because the site truncates long link text ("A Light in the ..."). Reading textContent would give you the shortened version.

Scoped queries like firstBook.$eval() are the cleanest way to read fields from a card. The selector h3 a only searches inside that one article.

Step 2: Get every match with page.$$()

page.$$() is the querySelectorAll() equivalent. It returns an array of handles, or an empty array when nothing matches.

const cards = await page.$$('article.product_pod');
console.log(`Found ${cards.length} books`); // 20 per page on this site

for (const card of cards) {
  const price = await card.$eval('.price_color', el => el.textContent);
  console.log(price);
}

This works, but look at the loop. Each card.$eval() is a separate message from Node to Chrome and back over the DevTools Protocol, so 20 cards means 20 round trips.

For 20 elements you won't notice. For a results page with a few thousand rows, you will. Step 3 collapses the whole thing into one call.

Step 3: Extract data with $eval and $$eval

page.$$eval() finds every match and runs your callback on all of them inside the browser. Only the return value crosses back to Node.

const books = await page.$$eval('article.product_pod', cards =>
  cards.map(card => ({
    title: card.querySelector('h3 a').getAttribute('title'),
    price: card.querySelector('.price_color').textContent.trim(),
    inStock: card.querySelector('.availability').textContent.includes('In stock'),
  }))
);

console.log(books.length, books[0]);

The callback runs in Chrome, so card is a real DOM node and querySelector works as usual.

What comes back is a plain array of objects with nothing tied to the live page.

That last point is the reason I reach for $$eval first when scraping. Plain data can't go stale when the page navigates, and it drops straight into JSON or a database.

Passing Node variables into the callback

The callback can't see variables from your Node scope. Pass them as extra arguments after the callback instead.

const maxPrice = 20;

const cheapPrices = await page.$$eval(
  '.price_color',
  (els, limit) => els
    .map(el => parseFloat(el.textContent.replace('£', '')))
    .filter(price => price < limit),
  maxPrice // arrives in the browser as `limit`
);

Arguments get serialized on the way in, so stick to strings, numbers, arrays, and plain objects. Functions and class instances won't survive the trip.

textContent vs. innerText

textContent returns every text node, including text hidden with CSS. innerText returns what a user would see, respects display: none, and forces a layout calculation to get there.

For scraping I default to textContent plus .trim(). I switch to innerText when a page hides duplicate labels for screen readers and I only want the visible one.

When $eval throws

Unlike page.$(), page.$eval() throws failed to find element matching selector on a miss. For optional fields, catch it and fall back:

const rating = await page
  .$eval('.star-rating', el => el.className.replace('star-rating ', ''))
  .catch(() => null); // missing element becomes null

Be aware this also swallows real errors inside your callback. For fields you expect on every page, let it throw so you hear about selector breakage.

Step 4: Wait for elements on JavaScript-rendered pages

None of the $ methods wait. They snapshot the DOM at the moment they run, and on a single-page app that moment often comes before the data arrives.

page.goto() resolves on the load event by default. A page that fetches its content after load will still be empty when your first query fires.

waitForSelector

page.waitForSelector() polls until a match appears, then returns its handle. Pass visible: true to skip elements that exist but are hidden.

await page.goto('https://quotes.toscrape.com/js/');

// Blocks until a quote is attached and visible (or 10s pass)
await page.waitForSelector('.quote', { visible: true, timeout: 10_000 });

const quotes = await page.$$eval('.quote .text', els =>
  els.map(el => el.textContent)
);

The /js/ version of quotes.toscrape.com renders its quotes with a script instead of plain HTML, which makes it a handy test bed for timing bugs.

Waiting for one quote doesn't prove all of them rendered. When a list streams in, wait for a condition on the count with page.waitForFunction() instead.

await page.waitForFunction(
  () => document.querySelectorAll('.quote').length >= 10
);

Locators

Locators are what the Puppeteer docs recommend for interacting with elements. A locator re-runs its query until it finds a match, then checks the element before acting.

Before a click, a locator makes sure the element is in the viewport, visible, enabled, and holding a stable bounding box across two animation frames. That last check prevents clicks landing mid-animation.

// Read a value with the same waiting behavior
const firstAuthor = await page
  .locator('.quote .author')
  .map(el => el.textContent)
  .wait();

// Click, and wait for the navigation the click triggers
await Promise.all([
  page.waitForNavigation(),
  page.locator('li.next > a').click(),
]);

.map().wait() gives you locator waiting with $eval-style output. When you need a raw handle for something locators can't do, .waitHandle() returns one.

Locators time out after 30 seconds by default, inherited from the page. Override per call with .setTimeout(5000) when a slow failure costs more than a fast one.

Get elements in Puppeteer by class, ID, and attribute

Everything in this section is plain CSS. If it works in document.querySelectorAll() in DevTools, it works in Puppeteer.

By class

Prefix the class with a dot. Chain classes with no space to require both, and use a space to mean "descendant of".

// Element with BOTH classes: <p class="instock availability">
const stock = await page.$('.instock.availability');

// .price_color anywhere inside a .product_price container
const price = await page.$('.product_price .price_color');

Class names are the most fragile hook on modern sites. CSS modules and utility-class builds produce names like css-1x2y3z that change on every deploy.

By ID

Prefix the ID with #. The quotes.toscrape.com login form uses clean IDs, and it accepts any credentials, which makes it good practice.

await page.goto('https://quotes.toscrape.com/login');

await page.locator('#username').fill('demo');
await page.locator('#password').fill('demo');

await Promise.all([
  page.waitForNavigation(),
  page.locator('input[type="submit"]').click(),
]);

IDs tend to outlive class names because developers add them as hooks for labels and scripts. Reload the page twice before trusting one, though; some frameworks generate IDs at render time.

Watch for IDs that aren't valid CSS. An ID starting with a digit, or containing a dot or colon, breaks the # syntax: #user.name reads as ID user plus class name.

The fix is an attribute selector, which treats the value as a plain string:

const field = await page.$('[id="user.name"]');
const generated = await page.$('[id=":r1:"]'); // colon-style framework IDs

By attribute

Attribute selectors cover everything classes and IDs miss. data-testid and name attributes are the most stable hooks you'll find on most sites.

const email = await page.$('input[name="email"]');
const testHook = await page.$('[data-testid="checkout-button"]');

// Partial matches: contains, starts with, ends with
const catalogueLinks = await page.$$('a[href*="catalogue"]');
const secureLinks = await page.$$('a[href^="https://"]');
const pdfLinks = await page.$$('a[href$=".pdf"]');

For more selector patterns in form work, see our guide to automating forms with Puppeteer.

Get elements by text

Most tutorials still loop through $$eval results and compare textContent. Puppeteer has had a built-in text selector for a while, and it works in every method that takes a selector.

await page.goto('https://quotes.toscrape.com/');

// ::-p-text() matches the deepest element containing the text
await Promise.all([
  page.waitForNavigation(),
  page.locator('li.next ::-p-text(Next)').click(),
]);

::-p-text() matches "minimal" elements: the deepest node containing your text, not every ancestor that also contains it. Scoping it with li.next keeps a stray "Next" elsewhere on the page from matching first.

It's a substring match. ::-p-text(Log) would match both "Login" and "Logout", so use the full visible string.

Parentheses and quotes inside the text need escaping, because the selector parser treats them as syntax:

await page.locator('::-p-text(Checkout \\(2 items\\))').click();

When you need an exact match, add a filter to a locator. The callback runs in the browser, like $$eval:

await page
  .locator('button')
  .filter(button => button.textContent.trim() === 'Add to basket')
  .click();

Get elements by XPath now that page.$x() is gone

page.$x() and page.waitForXPath() were removed in Puppeteer 22.0.0 in February 2024. Code that still calls them fails with TypeError: page.$x is not a function.

The replacement is the ::-p-xpath() selector, which works in $, $$, $eval, $$eval, waitForSelector, and locators.

await page.goto('https://quotes.toscrape.com/');

// Old, throws on v22+:
// const [el] = await page.$x('//small[@class="author"]');

const authors = await page.$$eval(
  '::-p-xpath(//small[@class="author"])',
  els => els.map(el => el.textContent)
);

I only use XPath for one thing CSS can't do: walking up or sideways from an element. Find the author, then climb to the quote card that contains it:

const einsteinCard = await page.$(
  '::-p-xpath(//small[text()="Albert Einstein"]/ancestor::div[@class="quote"])'
);

const quote = await einsteinCard?.$eval('.text', el => el.textContent);

The optional chaining (?.) handles the null case in one expression. For anything that only walks downward, CSS is shorter and easier to read six months later.

Get elements by ARIA role and name

::-p-aria() queries the browser's accessibility tree. It matches on the computed accessible name and role, so it keeps working when class names and DOM structure change underneath.

// Any element whose accessible name is "Login"
const login = await page.$('::-p-aria(Login)');

// Narrowed by role, so a "Login" heading won't match
const loginLink = await page.$('::-p-aria([name="Login"][role="link"])');

On sites built with hashed class names, this is the selector I trust most.

A redesign can rename every class, but the "Login" link still has to be a link named "Login" for screen readers to work.

Get elements inside shadow DOM and iframes

page.$() searches the main document only. It won't look inside shadow roots or child frames, and a miss there looks identical to a typo.

Shadow DOM: the >>> and >>>> combinators

>>> searches every open shadow root below a host, at any depth. >>>> only looks inside the host's immediate shadow root.

This example builds a small web component with page.setContent(), so you can run it without a target site:

await page.setContent(`
  <price-widget></price-widget>
  <script>
    const host = document.querySelector('price-widget');
    const root = host.attachShadow({ mode: 'open' });
    root.innerHTML = '<span class="amount">$19.99</span>';
  </script>
`);

const plain = await page.$('price-widget .amount'); // null: CSS stops at the shadow boundary
const amount = await page.$eval('price-widget >>> .amount', el => el.textContent);

The combinators only reach open shadow roots. A component created with mode: 'closed' stays sealed to selectors.

Iframes: get the frame first

Each iframe has its own document. Grab the <iframe> element, turn it into a Frame with contentFrame(), and query that.

await page.setContent(`
  <iframe name="checkout" srcdoc="<button id='pay'>Pay now</button>"></iframe>
`);

const frameElement = await page.waitForSelector('iframe[name="checkout"]');
const frame = await frameElement.contentFrame();

// Frames have the same $, $$, $eval, locator methods as pages
await frame.waitForSelector('#pay');
const label = await frame.$eval('#pay', el => el.textContent);

Payment forms and embedded widgets are the usual culprits here. If DevTools shows your element but Puppeteer can't find it, check the Elements panel for an #document node above it.

Full example: scrape a paginated catalog without stale handles

This is where the handle lifecycle stops being theory. An ElementHandle points at one node in one document, and navigation throws that document away.

Hold a handle across a click that loads a new page and it's dead, even if an identical element now sits in the same spot.

The pattern below avoids that by keeping handles short-lived.

First, a function that returns plain data from whatever page is loaded:

async function extractBooks(page) {
  await page.waitForSelector('article.product_pod');

  return page.$$eval('article.product_pod', cards =>
    cards.map(card => ({
      title: card.querySelector('h3 a').getAttribute('title'),
      price: card.querySelector('.price_color').textContent.trim(),
    }))
  );
}

Then the loop. Every query runs fresh on every page, and the only handle (next) is used once, immediately:

await page.goto('https://books.toscrape.com/catalogue/page-1.html');
const results = [];

for (let pageNum = 1; pageNum <= 3; pageNum++) {
  results.push(...(await extractBooks(page)));

  const next = await page.$('li.next > a');
  if (!next) break; // last page has no "next" link

  await Promise.all([page.waitForNavigation(), next.click()]);
}

console.log(`Scraped ${results.length} books`); // 60 for 3 pages

Promise.all starts listening for the navigation before the click fires. Awaiting the click first and the navigation second can miss a fast page load and hang until timeout.

The rule I follow: extract data, not handles. Anything that needs to outlive the current page should be a string or an object before the page changes.

Our Puppeteer web scraping guide builds this into a full crawler with storage and retries.

Troubleshooting

"TypeError: Cannot read properties of null (reading 'click')"

Why: page.$() found nothing and returned null, and the next line called a method on it.

Fix: test the selector in the DevTools console with document.querySelectorAll('your-selector'). If it matches there, the element renders late; replace page.$() with page.locator() or add page.waitForSelector().

"Error: failed to find element matching selector"

Why: $eval throws on a miss, while $ returns null. Scripts that mix the two get inconsistent failures.

Fix: wait for the element before calling $eval, or add .catch(() => null) for fields that are allowed to be missing.

"TimeoutError: Waiting for selector ... failed"

The full message looks like this:

TimeoutError: Waiting for selector `.price` failed: Waiting failed: 30000ms exceeded

Why: the element never appeared. Usual causes are a wrong selector, an element inside an iframe or shadow root, or a different page from the one you expected.

Fix: grab a screenshot at the failure point with await page.screenshot({ path: 'debug.png' }) and check what Chrome received.

Our Puppeteer screenshots guide covers full-page and element captures for this kind of debugging.

If the screenshot shows a CAPTCHA or an "access denied" page, your selector is fine and the site is blocking you. Start with why your web scraper gets blocked.

For IP-based blocks, moving datacenter traffic to residential or ISP proxies (we run both at Roundproxies) is the usual fix.

"TypeError: page.$x is not a function"

Why: you're on Puppeteer 22 or newer, where $x() no longer exists.

Fix: move the expression into ::-p-xpath(). page.$x('//h2') becomes page.$$('::-p-xpath(//h2)'), and page.waitForXPath() becomes page.waitForSelector('::-p-xpath(...)').

"Execution context was destroyed" or "Node is detached from document"

Why: both mean the element you're pointing at no longer exists.

The first fires when a query runs mid-navigation. The second fires when a framework re-renders and replaces the node your handle pointed to.

Fix: wrap navigating clicks in Promise.all with page.waitForNavigation(), and re-query after any navigation or re-render. Locators help here, since they resolve the selector again on each attempt.

FAQ

What does page.$ return if no element is found?

page.$() returns null and doesn't throw. page.$$() returns an empty array. page.$eval() is the exception: it throws failed to find element matching selector.

How do I get the text of an element in Puppeteer?

Use await page.$eval('selector', el => el.textContent) for one element, or page.$$eval() with .map() for many. On pages that render late, use page.locator('selector').map(el => el.textContent).wait() so the read waits for the element.

What replaced page.$x() in Puppeteer?

The ::-p-xpath() pseudo-element replaced it. Pass '::-p-xpath(//your/expression)' to any selector method, including page.$$(), page.waitForSelector(), and page.locator().

What's the difference between page.$ and page.locator?

page.$() takes one snapshot of the DOM and returns a handle or null. page.locator() retries until the element exists and is visible, enabled, and stable.

Use locators for interaction and $ when you need a handle right now.

Wrapping up

The mental model to keep: $ and $$ get handles, $eval and $$eval get data, and locators get handles that wait.

Pick based on whether you'll act on the element or save what's in it.

For new code, start with $$eval for extraction and locators for clicks, and use the P-selectors (::-p-text, ::-p-aria, ::-p-xpath) before writing any custom page.evaluate() matching logic.

When your scraper starts getting fingerprinted rather than failing on selectors, the next read is our guide to Puppeteer stealth.