How to extract fields with CSS selectors
Fetching a whole page is the easy part. What you usually want is five values from it, as JSON, without writing a parser. This guide shows how to extract fields with CSS selectors using POST /v1/extract: text, attributes, lists, and groups that give you one clean object per product. All examples run against books.toscrape.com, a shop built for practice, and every response is real output.
The basic request
You send a url and a fields object. Each key in fields is a name you choose; each value is a selector. You get back a data object with the same keys.
curl https://scrape.land/v1/extract \
-H "X-Api-Key: YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://books.toscrape.com/",
"fields": {
"heading": "h1",
"first_title": "article.product_pod h3 a@title",
"first_price": "article.product_pod .price_color",
"next_page": "li.next a@href"
}}'{
"data": {
"first_price": "£51.77",
"first_title": "A Light in the Attic",
"heading": "All products",
"next_page": "catalogue/page-2.html"
},
"status": 200,
"url": "https://books.toscrape.com/"
}Two rules cover most of it:
- A plain selector returns the trimmed text of the first match.
"h1"gave"All products". selector@attrreturns an attribute instead. The visible link text on this shop is cut short with "...", soh3 a@titlereads the full title from thetitleattribute.li.next a@hrefreads the link target.
Attribute values come back exactly as they are in the HTML, so next_page is relative. Resolve it against the page URL before you request it (urllib.parse.urljoin in Python, new URL(href, base) in JavaScript).
The object form: attributes and lists
Instead of a string, a field can be an object with css plus optional keys: attr to read an attribute, and all to return every match as an array. Here is a single product page, with the breadcrumb trail read as a list:
curl https://scrape.land/v1/extract \
-H "X-Api-Key: YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html",
"fields": {
"title": "h1",
"price": "p.price_color",
"stock": "p.instock.availability",
"image": "#product_gallery img@src",
"breadcrumbs": {"css": "ul.breadcrumb li", "all": true}
}}'{
"data": {
"breadcrumbs": ["Home", "Books", "Poetry", "A Light in the Attic"],
"image": "../../media/cache/fe/72/fe72f0532301ec28892ae79a629a293c.jpg",
"price": "£51.77",
"stock": "In stock (22 available)",
"title": "A Light in the Attic"
},
"status": 200,
"url": "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html"
}Whitespace around text is trimmed for you: the stock line is surrounded by newlines and an icon in the HTML, and comes back as one clean string.
Groups: one object per item
A category page has twenty books. You could ask for every title with all and every price with all, and zip the two arrays yourself. That works until one book is missing a price, at which point every price after it belongs to the wrong book.
A group avoids that. It is an object with css for the repeating container and its own fields, which are read inside each container. You get a list of objects, and a value missing from one item is null in that item only.
curl https://scrape.land/v1/extract \
-H "X-Api-Key: YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://books.toscrape.com/",
"fields": {
"books": {
"css": "article.product_pod",
"fields": {
"title": {"css": "h3 a", "attr": "title"},
"price": ".price_color",
"rating": {"css": "p.star-rating", "attr": "class"},
"url": "h3 a@href"
}
}
}}'{
"data": {
"books": [
{"price": "£51.77", "rating": "star-rating Three", "title": "A Light in the Attic", "url": "catalogue/a-light-in-the-attic_1000/index.html"},
{"price": "£53.74", "rating": "star-rating One", "title": "Tipping the Velvet", "url": "catalogue/tipping-the-velvet_999/index.html"},
{"price": "£50.10", "rating": "star-rating One", "title": "Soumission", "url": "catalogue/soumission_998/index.html"}
]
},
"status": 200,
"url": "https://books.toscrape.com/"
}We trimmed the list to three of the twenty rows. Note the rating: this shop encodes it in a class name (star-rating Three), which is exactly the kind of data a selector reaching into attributes handles well. Inside a group, "." selects the container itself, if you need an attribute of the container.
The same request from Python, turning the rows into clean records:
from urllib.parse import urljoin
import requests
PAGE = "https://books.toscrape.com/"
WORDS = {"One": 1, "Two": 2, "Three": 3, "Four": 4, "Five": 5}
r = requests.post(
"https://scrape.land/v1/extract",
headers={"X-Api-Key": "YOUR_KEY"},
json={"url": PAGE, "fields": {"books": {
"css": "article.product_pod",
"fields": {
"title": {"css": "h3 a", "attr": "title"},
"price": ".price_color",
"rating": {"css": "p.star-rating", "attr": "class"},
"url": "h3 a@href",
},
}}},
timeout=60,
)
r.raise_for_status()
for b in r.json()["data"]["books"]:
stars = WORDS.get((b["rating"] or "").replace("star-rating", "").strip())
print(b["title"], b["price"], stars, urljoin(PAGE, b["url"] or ""))Its first line of output: A Light in the Attic £51.77 3 https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html, followed by the other nineteen books.
Null values and field_errors
A field is null for one of two reasons: the page really has no such element, or your selector could not be used at all. The API tells you which. Here we asked for a .sale-badge the shop does not have, and used :is(), which the selector engine does not support:
{
"data": {"bad": null, "heading": "All products", "sale_badge": null},
"field_errors": {
"bad": "selector did not compile (unknown pseudoclass or pseudoelement :is) … this field is null because the selector could not be used, not because the page lacks the element"
},
"status": 200,
"url": "https://books.toscrape.com/"
}sale_badge is null with no error, so it is genuinely not on the page. bad is named in field_errors, so the selector is the problem. The request still succeeded and heading is unaffected: one broken selector never sinks the rest. In code, check for field_errors on every response and alert on it; it is the early warning that a selector needs fixing.
Selector tips
- Everything in CSS3 works: descendant and child combinators,
:nth-child(),:not(), and attribute selectors such as[href^="/catalogue"]. Newer selectors like:is(),:where()and relative:has()do not yet. - Matching on text:
:contains("In stock")matches an element whose text contains a string. It is an extension of this API, not standard CSS, so it will not work in your browser's devtools. - Copying selectors from devtools: "Copy selector" often gives something brittle like
#default > div > div > div. Prefer short selectors built on class names that describe the content. - Empty results on a page that looks fine in your browser: the content may be built by JavaScript. Check with a plain
/v1/fetch; if the values are not in the HTML, you need"render": true.
What it costs
A CSS extraction costs the same as a plain fetch: 1 request unit per page, on every plan including Free. Failed and blocked requests are not billed. See pricing.
Next steps
When CSS cannot express what you need, such as "the cell next to the heading that says UPC", read how to extract data with XPath selectors. To follow the next-page link across a whole category, see scraping a paginated catalog. The full selector reference is in the docs.