Skip to content
All posts

How to extract data with XPath selectors

Getting started6 min read
How to extract data with XPath selectors: the post's first code sample

CSS selectors cover most scraping jobs, but they only walk down the page and they cannot look at text. XPath can do both: find the table cell next to a heading that says "UPC", or start at an author's name and climb back up to their quote. This guide shows how to extract data with XPath selectors in POST /v1/extract, when XPath is the better tool, and the few rules to know. Every response is real output.

How XPath fits into a request

XPath goes in the same fields object as CSS, and you can mix the two in one request. There are two ways to write it:

An expression can end in text() or @attribute, or select elements, in which case you get their trimmed text. As with CSS, a string field returns the first match and a field that matches nothing is null.

Match by text: the cell next to a label

Product pages often have a specification table: a heading cell with a label, and a value cell next to it. The rows have no classes, and their order can change between products. CSS cannot say "the cell next to the one that reads UPC". XPath can:

curl
curl https://scrape.land/v1/extract \
  -H "X-Api-Key: YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html",
       "fields": {
         "title":    "//h1/text()",
         "upc":      "//th[text()=\"UPC\"]/following-sibling::td",
         "tax":      "//th[.=\"Tax\"]/following-sibling::td",
         "reviews":  "//tr[th=\"Number of reviews\"]/td",
         "category": "//ul[@class=\"breadcrumb\"]/li[last()-1]/a",
         "rating":   "//p[contains(@class,\"star-rating\")]/@class"
       }}'
JSON
{
  "data": {
    "category": "Poetry",
    "rating": "star-rating Three",
    "reviews": "0",
    "tax": "£0.00",
    "title": "A Light in the Attic",
    "upc": "a897fe39b1053632"
  },
  "status": 200,
  "url": "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html"
}

What each expression does:

Inside the JSON body, double quotes in the expression must be escaped as \". In Python or JavaScript, use single quotes inside the XPath and let the JSON library handle the rest.

Climb up: from a name to its quote

CSS selectors only go down the tree. When the thing you can identify (an author's name, a tag, a "Sold out" label) sits inside the block you want, XPath lets you find it first and then climb with .. or ancestor::. On quotes.toscrape.com:

curl
curl https://scrape.land/v1/extract \
  -H "X-Api-Key: YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://quotes.toscrape.com/",
       "fields": {
         "einstein_quotes": {
           "xpath": "//small[@class=\"author\"][.=\"Albert Einstein\"]/../../span[@class=\"text\"]",
           "all": true
         },
         "author_of_world_quote": "//span[@class=\"text\"][contains(., \"world\")]/../span/small",
         "love_tagged": {
           "xpath": "//a[@class=\"tag\"][.=\"love\"]/ancestor::div[@class=\"quote\"]/span/small",
           "all": true
         },
         "next": "//li[@class=\"next\"]/a/@href"
       }}'
JSON
{
  "data": {
    "author_of_world_quote": "Albert Einstein",
    "einstein_quotes": [
      "“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”",
      "“There are only two ways to live your life. One is as though nothing is a miracle. The other is as though everything is a miracle.”",
      "“Try not to become a man of success. Rather become a man of value.”"
    ],
    "love_tagged": ["André Gide"],
    "next": "/page/2/"
  },
  "status": 200,
  "url": "https://quotes.toscrape.com/"
}

The author's name is inside a small, inside a span, next to the quote text. /../.. climbs two levels to the quote block, and span[@class="text"] comes back down to the quote. ancestor::div[@class="quote"] does the same without counting levels, which is sturdier when the nesting varies. The page has three Einstein quotes, and only one quote tagged "love".

In Python

Single quotes inside the expression keep the code readable:

Python
import requests

def quotes_by(author, page="https://quotes.toscrape.com/"):
    xp = f"//small[@class='author'][.='{author}']/../../span[@class='text']"
    r = requests.post(
        "https://scrape.land/v1/extract",
        headers={"X-Api-Key": "YOUR_KEY"},
        json={"url": page, "fields": {"quotes": {"xpath": xp, "all": True}}},
        timeout=60,
    )
    r.raise_for_status()
    return r.json()["data"]["quotes"]

for q in quotes_by("Albert Einstein"):
    print(q)

If a name can contain an apostrophe, switch the quotes around it ([.="O'Brien"]), because XPath 1.0 has no escape character inside a string.

Rules worth knowing

CSS or XPath?

Use CSS by default: it is shorter, most developers read it at a glance, and it works inside groups. Reach for XPath when you need to match on text, move sideways (following-sibling), climb up (.., ancestor::), or count from the end (last()). This API also supports :contains() in CSS for simple text matching, so for "an element that contains this text" either works.

What it costs

XPath extraction costs the same as CSS and as a plain fetch: 1 request unit per page, on every plan including Free. Failed and blocked requests are not billed. See pricing.

Next steps

If you have not used CSS fields yet, start with how to extract fields with CSS selectors. For turning a specification table into rows, see extracting HTML tables into CSV. The selector reference is in the docs.

Start free with 1,000 requests Read the docs