Skip to content
All posts

How to scrape a product catalog with pagination

E-commerce and prices4 min read

Most catalogs split their products across numbered pages with a "next" link at the bottom. This tutorial reads one page of products and that page's next link in a single call, then loops in Python until the catalog runs out, and writes everything to a CSV. The example target is books.toscrape.com, a shop built for scraping practice.

The Books to Scrape catalog page, 1000 results showing 1 to 20, as captured by the scrape.land screenshot feature
The catalog as the API sees it, captured with "screenshot": true on POST /v1/fetch. 1,000 books, 20 per page.

Why the loop lives in your code

The API is stateless on purpose: one request fetches one page. It does not follow next links for you, and that is a feature. You decide how many pages to read, how fast, in what order, and what to do when a page fails. A loop that you own is also the easiest thing to resume: store the last URL you finished and start from there next time.

So the trick is to make each call return two things: the items on the page, and the URL of the next page. Then the loop is five lines.

Step 1: extract one page

Each book on the page is an article.product_pod. A nested field (a css plus its own fields) gives you one object per book, with every value read from inside that book's own card, so a missing price can never shift the rows around. The selector@attr shorthand reads an attribute instead of text, which is how you get the full title and the link.

curl
curl https://scrape.land/v1/extract \
  -H "X-Api-Key: YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://books.toscrape.com/catalogue/page-1.html",
       "fields": {
         "books": {
           "css": "article.product_pod",
           "fields": {"title": "h3 a@title", "price": ".price_color",
                      "stock": ".availability", "url": "h3 a@href"}
         },
         "next": "li.next a@href"
       }}'

This is what came back when we ran it (trimmed to three of the twenty books):

Terminal showing the curl request to /v1/extract and the JSON response with three books, their prices, stock and URLs, plus the next field set to page-2.html
The real response. next is page-2.html, and on the last page it is null.

Two details matter for the loop. The url and next values are exactly what the page's href attributes say, and on this site they are relative (page-2.html), so resolve them against the page URL. And on the last page there is no li.next, so next comes back null, which is your stop condition.

Step 2: loop and write a CSV

Python
import csv
from urllib.parse import urljoin

import requests

API = "https://scrape.land/v1/extract"
FIELDS = {
    "books": {
        "css": "article.product_pod",
        "fields": {
            "title": "h3 a@title",
            "price": ".price_color",
            "stock": ".availability",
            "url": "h3 a@href",
        },
    },
    "next": "li.next a@href",
}

url = "https://books.toscrape.com/catalogue/page-1.html"
rows, pages = [], 0
while url and pages < 3:  # raise the cap when you are happy with the output
    r = requests.post(API, headers={"X-Api-Key": "YOUR_KEY"},
                      json={"url": url, "fields": FIELDS}, timeout=60)
    r.raise_for_status()
    data = r.json()["data"]
    for b in data["books"]:
        b["url"] = urljoin(url, b["url"])  # hrefs on the page are relative
        rows.append(b)
    pages += 1
    url = urljoin(url, data["next"]) if data["next"] else None

with open("books.csv", "w", newline="", encoding="utf-8") as f:
    w = csv.DictWriter(f, fieldnames=["title", "price", "stock", "url"])
    w.writeheader()
    w.writerows(rows)
print(f"{len(rows)} books from {pages} pages, next page: {url}")

We ran exactly this script with a cap of three pages:

Terminal output: 60 books from 3 pages, next page page-4.html, followed by the first rows of books.csv with title, price and stock columns
Three calls, 60 rows, and the URL where the next run would continue.

Making it production-ready

What it costs

Each page is one plain fetch: 1 request unit. The three-page run above cost 3 units, and the full 50-page catalog would cost 50. No rendering is needed on this site. The Free plan includes 1,000 requests a month with no card, and you only pay for responses that land. See pricing.

Next steps

Selectors, nested fields and the @attr shorthand are documented in the docs. If the catalog loads more products as you scroll instead of using numbered pages, read how to scrape infinite-scroll pages. To track the prices over time, see monitoring prices on a schedule. Create a free account to run the script with your own key.

Start free with 1,000 requests Read the docs