Skip to content
All posts

How to extract fields with CSS selectors

Getting started6 min read
How to extract fields with CSS selectors: the post's first code sample

Fetching a whole page is the easy part. What you usually want is five values from it, as JSON, without writing a parser. This guide shows how to extract fields with CSS selectors using POST /v1/extract: text, attributes, lists, and groups that give you one clean object per product. All examples run against books.toscrape.com, a shop built for practice, and every response is real output.

The basic request

You send a url and a fields object. Each key in fields is a name you choose; each value is a selector. You get back a data object with the same keys.

curl
curl https://scrape.land/v1/extract \
  -H "X-Api-Key: YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://books.toscrape.com/",
       "fields": {
         "heading":     "h1",
         "first_title": "article.product_pod h3 a@title",
         "first_price": "article.product_pod .price_color",
         "next_page":   "li.next a@href"
       }}'
JSON
{
  "data": {
    "first_price": "£51.77",
    "first_title": "A Light in the Attic",
    "heading": "All products",
    "next_page": "catalogue/page-2.html"
  },
  "status": 200,
  "url": "https://books.toscrape.com/"
}

Two rules cover most of it:

Attribute values come back exactly as they are in the HTML, so next_page is relative. Resolve it against the page URL before you request it (urllib.parse.urljoin in Python, new URL(href, base) in JavaScript).

The object form: attributes and lists

Instead of a string, a field can be an object with css plus optional keys: attr to read an attribute, and all to return every match as an array. Here is a single product page, with the breadcrumb trail read as a list:

curl
curl https://scrape.land/v1/extract \
  -H "X-Api-Key: YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html",
       "fields": {
         "title":       "h1",
         "price":       "p.price_color",
         "stock":       "p.instock.availability",
         "image":       "#product_gallery img@src",
         "breadcrumbs": {"css": "ul.breadcrumb li", "all": true}
       }}'
JSON
{
  "data": {
    "breadcrumbs": ["Home", "Books", "Poetry", "A Light in the Attic"],
    "image": "../../media/cache/fe/72/fe72f0532301ec28892ae79a629a293c.jpg",
    "price": "£51.77",
    "stock": "In stock (22 available)",
    "title": "A Light in the Attic"
  },
  "status": 200,
  "url": "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html"
}

Whitespace around text is trimmed for you: the stock line is surrounded by newlines and an icon in the HTML, and comes back as one clean string.

Groups: one object per item

A category page has twenty books. You could ask for every title with all and every price with all, and zip the two arrays yourself. That works until one book is missing a price, at which point every price after it belongs to the wrong book.

A group avoids that. It is an object with css for the repeating container and its own fields, which are read inside each container. You get a list of objects, and a value missing from one item is null in that item only.

curl
curl https://scrape.land/v1/extract \
  -H "X-Api-Key: YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://books.toscrape.com/",
       "fields": {
         "books": {
           "css": "article.product_pod",
           "fields": {
             "title":  {"css": "h3 a", "attr": "title"},
             "price":  ".price_color",
             "rating": {"css": "p.star-rating", "attr": "class"},
             "url":    "h3 a@href"
           }
         }
       }}'
JSON
{
  "data": {
    "books": [
      {"price": "£51.77", "rating": "star-rating Three", "title": "A Light in the Attic", "url": "catalogue/a-light-in-the-attic_1000/index.html"},
      {"price": "£53.74", "rating": "star-rating One", "title": "Tipping the Velvet", "url": "catalogue/tipping-the-velvet_999/index.html"},
      {"price": "£50.10", "rating": "star-rating One", "title": "Soumission", "url": "catalogue/soumission_998/index.html"}
    ]
  },
  "status": 200,
  "url": "https://books.toscrape.com/"
}

We trimmed the list to three of the twenty rows. Note the rating: this shop encodes it in a class name (star-rating Three), which is exactly the kind of data a selector reaching into attributes handles well. Inside a group, "." selects the container itself, if you need an attribute of the container.

The same request from Python, turning the rows into clean records:

Python
from urllib.parse import urljoin
import requests

PAGE = "https://books.toscrape.com/"
WORDS = {"One": 1, "Two": 2, "Three": 3, "Four": 4, "Five": 5}

r = requests.post(
    "https://scrape.land/v1/extract",
    headers={"X-Api-Key": "YOUR_KEY"},
    json={"url": PAGE, "fields": {"books": {
        "css": "article.product_pod",
        "fields": {
            "title": {"css": "h3 a", "attr": "title"},
            "price": ".price_color",
            "rating": {"css": "p.star-rating", "attr": "class"},
            "url": "h3 a@href",
        },
    }}},
    timeout=60,
)
r.raise_for_status()
for b in r.json()["data"]["books"]:
    stars = WORDS.get((b["rating"] or "").replace("star-rating", "").strip())
    print(b["title"], b["price"], stars, urljoin(PAGE, b["url"] or ""))

Its first line of output: A Light in the Attic £51.77 3 https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html, followed by the other nineteen books.

Null values and field_errors

A field is null for one of two reasons: the page really has no such element, or your selector could not be used at all. The API tells you which. Here we asked for a .sale-badge the shop does not have, and used :is(), which the selector engine does not support:

JSON
{
  "data": {"bad": null, "heading": "All products", "sale_badge": null},
  "field_errors": {
    "bad": "selector did not compile (unknown pseudoclass or pseudoelement :is) … this field is null because the selector could not be used, not because the page lacks the element"
  },
  "status": 200,
  "url": "https://books.toscrape.com/"
}

sale_badge is null with no error, so it is genuinely not on the page. bad is named in field_errors, so the selector is the problem. The request still succeeded and heading is unaffected: one broken selector never sinks the rest. In code, check for field_errors on every response and alert on it; it is the early warning that a selector needs fixing.

Selector tips

What it costs

A CSS extraction costs the same as a plain fetch: 1 request unit per page, on every plan including Free. Failed and blocked requests are not billed. See pricing.

Next steps

When CSS cannot express what you need, such as "the cell next to the heading that says UPC", read how to extract data with XPath selectors. To follow the next-page link across a whole category, see scraping a paginated catalog. The full selector reference is in the docs.

Start free with 1,000 requests Read the docs