How to extract data with XPath selectors
CSS selectors cover most scraping jobs, but they only walk down the page and they cannot look at text. XPath can do both: find the table cell next to a heading that says "UPC", or start at an author's name and climb back up to their quote. This guide shows how to extract data with XPath selectors in POST /v1/extract, when XPath is the better tool, and the few rules to know. Every response is real output.
How XPath fits into a request
XPath goes in the same fields object as CSS, and you can mix the two in one request. There are two ways to write it:
- As a string starting with
/or(. Anything that starts that way is treated as XPath:"//h1/text()","//a/@href","(//span[@class='text'])[1]". - As an object with
xpathinstead ofcss, which lets you add"all": truefor every match, orattrto read an attribute of the matched element.
An expression can end in text() or @attribute, or select elements, in which case you get their trimmed text. As with CSS, a string field returns the first match and a field that matches nothing is null.
Match by text: the cell next to a label
Product pages often have a specification table: a heading cell with a label, and a value cell next to it. The rows have no classes, and their order can change between products. CSS cannot say "the cell next to the one that reads UPC". XPath can:
curl https://scrape.land/v1/extract \
-H "X-Api-Key: YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html",
"fields": {
"title": "//h1/text()",
"upc": "//th[text()=\"UPC\"]/following-sibling::td",
"tax": "//th[.=\"Tax\"]/following-sibling::td",
"reviews": "//tr[th=\"Number of reviews\"]/td",
"category": "//ul[@class=\"breadcrumb\"]/li[last()-1]/a",
"rating": "//p[contains(@class,\"star-rating\")]/@class"
}}'{
"data": {
"category": "Poetry",
"rating": "star-rating Three",
"reviews": "0",
"tax": "£0.00",
"title": "A Light in the Attic",
"upc": "a897fe39b1053632"
},
"status": 200,
"url": "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html"
}What each expression does:
//th[text()="UPC"]/following-sibling::tdfinds the heading cell whose text is exactly "UPC", then moves sideways to the nexttd. Thefollowing-siblingaxis is the part CSS has no equivalent for.//th[.="Tax"]is the same idea;.means "this element's full text", which also works when the text is split across child elements.//tr[th="Number of reviews"]/tdselects the row that contains such a heading, then its value cell.li[last()-1]counts from the end: the second-to-last breadcrumb is the category, however deep the trail is.contains(@class,"star-rating")matches one class among several, and/@classreturns the attribute.
Inside the JSON body, double quotes in the expression must be escaped as \". In Python or JavaScript, use single quotes inside the XPath and let the JSON library handle the rest.
Climb up: from a name to its quote
CSS selectors only go down the tree. When the thing you can identify (an author's name, a tag, a "Sold out" label) sits inside the block you want, XPath lets you find it first and then climb with .. or ancestor::. On quotes.toscrape.com:
curl https://scrape.land/v1/extract \
-H "X-Api-Key: YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://quotes.toscrape.com/",
"fields": {
"einstein_quotes": {
"xpath": "//small[@class=\"author\"][.=\"Albert Einstein\"]/../../span[@class=\"text\"]",
"all": true
},
"author_of_world_quote": "//span[@class=\"text\"][contains(., \"world\")]/../span/small",
"love_tagged": {
"xpath": "//a[@class=\"tag\"][.=\"love\"]/ancestor::div[@class=\"quote\"]/span/small",
"all": true
},
"next": "//li[@class=\"next\"]/a/@href"
}}'{
"data": {
"author_of_world_quote": "Albert Einstein",
"einstein_quotes": [
"“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”",
"“There are only two ways to live your life. One is as though nothing is a miracle. The other is as though everything is a miracle.”",
"“Try not to become a man of success. Rather become a man of value.”"
],
"love_tagged": ["André Gide"],
"next": "/page/2/"
},
"status": 200,
"url": "https://quotes.toscrape.com/"
}The author's name is inside a small, inside a span, next to the quote text. /../.. climbs two levels to the quote block, and span[@class="text"] comes back down to the quote. ancestor::div[@class="quote"] does the same without counting levels, which is sturdier when the nesting varies. The page has three Einstein quotes, and only one quote tagged "love".
In Python
Single quotes inside the expression keep the code readable:
import requests
def quotes_by(author, page="https://quotes.toscrape.com/"):
xp = f"//small[@class='author'][.='{author}']/../../span[@class='text']"
r = requests.post(
"https://scrape.land/v1/extract",
headers={"X-Api-Key": "YOUR_KEY"},
json={"url": page, "fields": {"quotes": {"xpath": xp, "all": True}}},
timeout=60,
)
r.raise_for_status()
return r.json()["data"]["quotes"]
for q in quotes_by("Albert Einstein"):
print(q)If a name can contain an apostrophe, switch the quotes around it ([.="O'Brien"]), because XPath 1.0 has no escape character inside a string.
Rules worth knowing
- Invalid XPath fails the request. A syntax error in an XPath field returns
400naming the field, before anything is fetched, so you find the mistake immediately. We gotfield "t": invalid xpath: expression must evaluate to a node-setfor an unfinished//h1[@. - The expression must select nodes. Functions that return a number or a string, such as
count(//div), are not supported as a field. A string that does not start with/or(is read as CSS, socount(//div)came backnullwith afield_errorsentry saying the selector did not compile. Count the items of an"all": truelist in your own code instead. - XPath is top-level only. A group's container must be selected with
css;{"xpath": ..., "fields": ...}is a400("a group must select its container with "css", not "xpath""). XPath fields inside a group also come backnullfor every row, because they would otherwise search the whole document and mix up rows. Use CSS inside groups, and XPath at the top level. - Write for the HTML, not the browser. XPath runs against the HTML the server sent. Browsers insert elements such as
tbodyinto tables that do not have them, so an XPath copied from devtools may include atbodystep that is not in the source. Start expressions with//and anchor on attributes and text rather than long absolute paths.
CSS or XPath?
Use CSS by default: it is shorter, most developers read it at a glance, and it works inside groups. Reach for XPath when you need to match on text, move sideways (following-sibling), climb up (.., ancestor::), or count from the end (last()). This API also supports :contains() in CSS for simple text matching, so for "an element that contains this text" either works.
What it costs
XPath extraction costs the same as CSS and as a plain fetch: 1 request unit per page, on every plan including Free. Failed and blocked requests are not billed. See pricing.
Next steps
If you have not used CSS fields yet, start with how to extract fields with CSS selectors. For turning a specification table into rows, see extracting HTML tables into CSV. The selector reference is in the docs.