Skip to content
All posts

How to extract product prices into a spreadsheet

E-commerce and prices5 min read
How to extract product prices into a spreadsheet: the post's first code sample

You have a category page or a list of product URLs, and you want a spreadsheet with one row per product: name, price, link. This guide does it two ways: with CSS selectors, which are exact and cheap, and with a plain-English prompt, which works on sites whose markup you do not want to study.

Option 1: CSS selectors

Open the page in your browser, right-click a product and choose Inspect. You are looking for the element that wraps one product (often something like .product-card or li.product) and, inside it, the elements that hold the name and the price.

Then send those selectors to POST /v1/extract as a group: an object with a css selector for the repeating container and its own fields, which are read inside each container. You get back a list of objects, one per product, so a product with a missing price is null in its own row instead of shifting every row below it.

curl
curl https://scrape.land/v1/extract \
  -H "X-Api-Key: YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://shop.example/category/desk-lamps",
       "fields": {
         "products": {
           "css": ".product-card",
           "fields": {
             "name":  "h2",
             "price": ".price",
             "url":   {"css": "a", "attr": "href"}
           }
         }
       }}'
JSON
{
  "url": "https://shop.example/category/desk-lamps",
  "status": 200,
  "data": {
    "products": [
      {"name": "Vintage desk lamp", "price": "$49.00", "url": "/item/42"},
      {"name": "Clamp lamp, black", "price": "$27.50", "url": "/item/43"},
      {"name": "Brass reading lamp", "price": null, "url": "/item/44"}
    ]
  }
}

The third product has no .price element (perhaps it is sold out), so its price is null. If a selector itself is invalid, the response adds a field_errors object naming the field, so you can tell "the page has no price" apart from "my selector is broken".

Into a CSV with Python

Python
import csv
import re
from urllib.parse import urljoin
import requests

PAGE = "https://shop.example/category/desk-lamps"
FIELDS = {
    "products": {
        "css": ".product-card",
        "fields": {"name": "h2", "price": ".price", "url": {"css": "a", "attr": "href"}},
    }
}

r = requests.post(
    "https://scrape.land/v1/extract",
    headers={"X-Api-Key": "YOUR_KEY"},
    json={"url": PAGE, "fields": FIELDS},
    timeout=60,
)
r.raise_for_status()
products = r.json()["data"]["products"] or []

def to_number(price):
    """'$1,049.00' -> 1049.0; None stays None."""
    if not price:
        return None
    digits = re.sub(r"[^0-9.]", "", price.replace(",", ""))
    return float(digits) if digits else None

with open("prices.csv", "w", newline="", encoding="utf-8") as f:
    w = csv.writer(f)
    w.writerow(["name", "price", "price_text", "url"])
    for p in products:
        w.writerow([p["name"], to_number(p["price"]), p["price"], urljoin(PAGE, p["url"] or "")])

Keep the original price text next to the parsed number. Currency formats vary (1.049,00 € is a thousand euros in much of Europe), and having the raw text makes a bad parse easy to spot. Open prices.csv in Excel, Numbers or Google Sheets.

If you have a list of product URLs rather than a category page, loop over them and send a flat field map ({"name": "h1", "price": ".price"}) for each. If the category spans several pages, fetch each page URL in turn: the API fetches exactly the URL you send and does not follow pagination for you.

Option 2: describe the fields in plain English

Selectors are the right tool when you scrape the same site repeatedly. When you have many different shops, writing selectors per site is the whole cost of the project. AI extraction takes a prompt instead of fields and reads the page for you. It is available from the Scale plan up.

curl
curl https://scrape.land/v1/extract \
  -H "X-Api-Key: YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://shop.example/category/desk-lamps",
       "prompt": "Every product on the page: name, price as a number, currency code, product URL",
       "model": "fast"}'
JSON
{
  "url": "https://shop.example/category/desk-lamps",
  "status": 200,
  "extracted_by": "ai",
  "model": "fast",
  "data": {
    "items": [
      {"name": "Vintage desk lamp", "price": 49.0, "currency": "USD", "url": "https://shop.example/item/42"},
      {"name": "Clamp lamp, black", "price": 27.5, "currency": "USD", "url": "https://shop.example/item/43"}
    ]
  }
}

When the answer is a list, it comes back under items. To make every shop return the same keys, send a schema with the prompt; the docs show the format. The Python above works unchanged if you replace the request body and read data["items"].

Which one to use

When prices look wrong

What it costs

A CSS extraction is 1 request unit per page on every plan. AI extraction (Scale plan and up) adds units per page: +2 on fast, +4 on smart, +25 on max. You only pay for responses that land. See pricing.

Next steps

To run this every day and catch price changes, see monitoring competitor prices on a schedule. If the prices only appear after JavaScript runs, read how to scrape a JavaScript-heavy page. Sign up free to get a key.

Start free with 1,000 requests Read the docs