Skip to content
All posts

How to crawl thousands of pages at 10 requests a second

Plans and scale6 min read
How to crawl thousands of pages at 10 requests a second: the post's first code sample

You have a list of a few thousand URLs and a plan that allows 10 requests per second, like the $19 Pool plan. Firing them all at once earns you a wall of 429 responses; sending them one by one wastes most of your limit waiting for slow pages. This guide builds a small Python crawler that runs at 10 requests a second, handles rate limits the way the docs describe, and picks up where it left off if you stop it.

Two numbers: rate and concurrency

The rate limit counts how many requests you start per second. How many you can start depends on how many are in flight, because each one waits for the target site to answer. If a page takes 1.5 seconds on average and you want 10 per second, you need about 15 requests open at once. With 4 open you would top out near 3 per second no matter how the limit is set.

So the crawler separates the two. A pool of worker threads keeps enough requests in flight, and a shared token bucket decides when each of them may start. The bucket refills at 10 tokens a second; a worker takes a token before every request and waits if none is left. Threads can be added freely without ever going over the rate.

The crawler

It reads URLs from urls.txt, extracts every book on each page with a CSS group field, and appends one JSON line per page to books.jsonl. It needs only requests.

Python
import json
import os
import threading
import time
from concurrent.futures import ThreadPoolExecutor

import requests

API = "https://scrape.land/v1/extract"
KEY = os.environ["SCRAPELAND_KEY"]
RATE = 10        # requests per second: your plan's limit
WORKERS = 16     # enough threads to keep the bucket busy
OUT = "books.jsonl"
FIELDS = {
    "books": {
        "css": "article.product_pod",
        "fields": {"title": "h3 a@title", "price": ".price_color", "url": "h3 a@href"},
    }
}


class TokenBucket:
    """Hands out at most `rate` tokens per second, shared by all threads."""

    def __init__(self, rate):
        self.rate = rate
        self.tokens = 1.0
        self.last = time.monotonic()
        self.lock = threading.Lock()

    def take(self):
        while True:
            with self.lock:
                now = time.monotonic()
                self.tokens = min(1.0, self.tokens + (now - self.last) * self.rate)
                self.last = now
                if self.tokens >= 1:
                    self.tokens -= 1
                    return
                wait = (1 - self.tokens) / self.rate
            time.sleep(wait)


bucket = TokenBucket(RATE)
write_lock = threading.Lock()
stats = {"ok": 0, "failed": 0, "429": 0}


def fetch(url):
    for attempt in range(6):
        bucket.take()
        try:
            r = requests.post(API, headers={"X-Api-Key": KEY},
                              json={"url": url, "fields": FIELDS}, timeout=90)
        except requests.RequestException:
            time.sleep(2 ** attempt)
            continue
        if r.status_code == 429:
            stats["429"] += 1
            time.sleep(float(r.headers.get("Retry-After", "1")))
            continue
        if r.status_code in (502, 503, 504):
            time.sleep(float(r.headers.get("Retry-After", 2 ** attempt)))
            continue
        if r.status_code == 402:
            raise SystemExit("402: quota or balance exhausted, stopping")
        body = r.json()
        with write_lock:
            with open(OUT, "a", encoding="utf-8") as f:
                row = {**body, "url": url, "http": r.status_code}
                f.write(json.dumps(row, ensure_ascii=False) + "\n")
            stats["ok" if r.ok else "failed"] += 1
        return
    stats["failed"] += 1


def main():
    done = set()
    if os.path.exists(OUT):
        with open(OUT, encoding="utf-8") as f:
            done = {json.loads(line)["url"] for line in f if line.strip()}
    with open("urls.txt", encoding="utf-8") as f:
        todo = [u.strip() for u in f if u.strip() and u.strip() not in done]
    print(f"{len(done)} already done, {len(todo)} to go")

    start = time.monotonic()
    with ThreadPoolExecutor(WORKERS) as pool:
        list(pool.map(fetch, todo))
    secs = time.monotonic() - start
    print(f"{stats['ok']} saved, {stats['failed']} failed, "
          f"{stats['429']} rate-limited retries, {secs:.1f}s")


if __name__ == "__main__":
    main()

How it handles errors

Note that http in each row is the status of our API, while the status field inside the body is what the target site answered. A 200 from us with a target status of 404 means the page really does not exist.

Resuming after a stop

Every finished page is appended to books.jsonl as soon as it arrives, under a lock so two threads never interleave a line. On start, the crawler reads the file, collects every URL already in it, and skips those. If you press Ctrl+C, lose your connection or hit a 402, run it again and it continues with what is left. To retry the failures, delete their lines from the file first.

Order is not preserved: threads finish in whatever order the pages answer. Every row carries its url, so sort afterwards if you need to.

A real run

We ran it against the first 20 catalogue pages of books.toscrape.com, a site built for scraping practice. The URL list is one line of shell:

shell
for i in $(seq 1 20); do echo "https://books.toscrape.com/catalogue/page-$i.html"; done > urls.txt
export SCRAPELAND_KEY=YOUR_KEY
python crawl.py
shell
0 already done, 20 to go
20 saved, 0 failed, 0 rate-limited retries, 26.3s

Then we added five more pages to urls.txt and ran it again. Only the new ones were fetched:

shell
20 already done, 5 to go
5 saved, 0 failed, 0 rate-limited retries, 12.1s

The 25 lines held 500 books. One line, cut to its first two books:

JSON
{"data": {"books": [{"price": "£12.84", "title": "In Her Wake", "url": "in-her-wake_980/index.html"},
 {"price": "£37.32", "title": "How Music Works", "url": "how-music-works_979/index.html"}, …]},
 "status": 200, "url": "https://books.toscrape.com/catalogue/page-2.html", "http": 200}

This test ran on the Free plan, whose exits are a high-latency public pool, so the time is set by how long each page took to come back rather than by the rate limit. A batch of 25 URLs is too small to reach 10 per second anyway: most of the run is waiting for the last few pages.

Scaling it up

What it costs

A CSS extraction is 1 request unit per page on metered plans, and only delivered pages are billed; a 429 is never billed. On the Pool plan there is no request count at all, only the 10 per second limit this crawler is built around. See pricing.

Next steps

Read what the Pool plan is for to see whether unmetered fits your crawl. The error codes and rate limits are listed in the docs. Sign up free to get a key and run the crawler on a few pages.

Start free with 1,000 requests Read the docs