How to crawl thousands of pages at 10 requests a second
You have a list of a few thousand URLs and a plan that allows 10 requests per second, like the $19 Pool plan. Firing them all at once earns you a wall of 429 responses; sending them one by one wastes most of your limit waiting for slow pages. This guide builds a small Python crawler that runs at 10 requests a second, handles rate limits the way the docs describe, and picks up where it left off if you stop it.
Two numbers: rate and concurrency
The rate limit counts how many requests you start per second. How many you can start depends on how many are in flight, because each one waits for the target site to answer. If a page takes 1.5 seconds on average and you want 10 per second, you need about 15 requests open at once. With 4 open you would top out near 3 per second no matter how the limit is set.
So the crawler separates the two. A pool of worker threads keeps enough requests in flight, and a shared token bucket decides when each of them may start. The bucket refills at 10 tokens a second; a worker takes a token before every request and waits if none is left. Threads can be added freely without ever going over the rate.
The crawler
It reads URLs from urls.txt, extracts every book on each page with a CSS group field, and appends one JSON line per page to books.jsonl. It needs only requests.
import json
import os
import threading
import time
from concurrent.futures import ThreadPoolExecutor
import requests
API = "https://scrape.land/v1/extract"
KEY = os.environ["SCRAPELAND_KEY"]
RATE = 10 # requests per second: your plan's limit
WORKERS = 16 # enough threads to keep the bucket busy
OUT = "books.jsonl"
FIELDS = {
"books": {
"css": "article.product_pod",
"fields": {"title": "h3 a@title", "price": ".price_color", "url": "h3 a@href"},
}
}
class TokenBucket:
"""Hands out at most `rate` tokens per second, shared by all threads."""
def __init__(self, rate):
self.rate = rate
self.tokens = 1.0
self.last = time.monotonic()
self.lock = threading.Lock()
def take(self):
while True:
with self.lock:
now = time.monotonic()
self.tokens = min(1.0, self.tokens + (now - self.last) * self.rate)
self.last = now
if self.tokens >= 1:
self.tokens -= 1
return
wait = (1 - self.tokens) / self.rate
time.sleep(wait)
bucket = TokenBucket(RATE)
write_lock = threading.Lock()
stats = {"ok": 0, "failed": 0, "429": 0}
def fetch(url):
for attempt in range(6):
bucket.take()
try:
r = requests.post(API, headers={"X-Api-Key": KEY},
json={"url": url, "fields": FIELDS}, timeout=90)
except requests.RequestException:
time.sleep(2 ** attempt)
continue
if r.status_code == 429:
stats["429"] += 1
time.sleep(float(r.headers.get("Retry-After", "1")))
continue
if r.status_code in (502, 503, 504):
time.sleep(float(r.headers.get("Retry-After", 2 ** attempt)))
continue
if r.status_code == 402:
raise SystemExit("402: quota or balance exhausted, stopping")
body = r.json()
with write_lock:
with open(OUT, "a", encoding="utf-8") as f:
row = {**body, "url": url, "http": r.status_code}
f.write(json.dumps(row, ensure_ascii=False) + "\n")
stats["ok" if r.ok else "failed"] += 1
return
stats["failed"] += 1
def main():
done = set()
if os.path.exists(OUT):
with open(OUT, encoding="utf-8") as f:
done = {json.loads(line)["url"] for line in f if line.strip()}
with open("urls.txt", encoding="utf-8") as f:
todo = [u.strip() for u in f if u.strip() and u.strip() not in done]
print(f"{len(done)} already done, {len(todo)} to go")
start = time.monotonic()
with ThreadPoolExecutor(WORKERS) as pool:
list(pool.map(fetch, todo))
secs = time.monotonic() - start
print(f"{stats['ok']} saved, {stats['failed']} failed, "
f"{stats['429']} rate-limited retries, {secs:.1f}s")
if __name__ == "__main__":
main()How it handles errors
429from the API means you went over your plan's rate. The docs say to wait for theRetry-Afterheader and continue; nothing is lost and the refused request is not billed. The worker sleeps that long and sends the same URL again. With the bucket in place you should rarely see one, but another script on the same account shares the limit, because it is enforced per account across all your keys.502,503and504are temporary on our side or the target's, so they are retried with a growing pause. A503carries its ownRetry-After, which the code respects.402means your quota or prepaid balance is used up. Retrying will not help, so the crawler stops.- Anything else (a
400for a bad URL, say) is written to the file with its status so you can look at it later, and the URL counts as done.
Note that http in each row is the status of our API, while the status field inside the body is what the target site answered. A 200 from us with a target status of 404 means the page really does not exist.
Resuming after a stop
Every finished page is appended to books.jsonl as soon as it arrives, under a lock so two threads never interleave a line. On start, the crawler reads the file, collects every URL already in it, and skips those. If you press Ctrl+C, lose your connection or hit a 402, run it again and it continues with what is left. To retry the failures, delete their lines from the file first.
Order is not preserved: threads finish in whatever order the pages answer. Every row carries its url, so sort afterwards if you need to.
A real run
We ran it against the first 20 catalogue pages of books.toscrape.com, a site built for scraping practice. The URL list is one line of shell:
for i in $(seq 1 20); do echo "https://books.toscrape.com/catalogue/page-$i.html"; done > urls.txt
export SCRAPELAND_KEY=YOUR_KEY
python crawl.py0 already done, 20 to go
20 saved, 0 failed, 0 rate-limited retries, 26.3sThen we added five more pages to urls.txt and ran it again. Only the new ones were fetched:
20 already done, 5 to go
5 saved, 0 failed, 0 rate-limited retries, 12.1sThe 25 lines held 500 books. One line, cut to its first two books:
{"data": {"books": [{"price": "£12.84", "title": "In Her Wake", "url": "in-her-wake_980/index.html"},
{"price": "£37.32", "title": "How Music Works", "url": "how-music-works_979/index.html"}, …]},
"status": 200, "url": "https://books.toscrape.com/catalogue/page-2.html", "http": 200}This test ran on the Free plan, whose exits are a high-latency public pool, so the time is set by how long each page took to come back rather than by the rate limit. A batch of 25 URLs is too small to reach 10 per second anyway: most of the run is waiting for the last few pages.
Scaling it up
- Thousands of URLs. At 10 per second, 5,000 pages take a little over 8 minutes of request starts, plus the tail of the slowest responses.
- Tuning
WORKERS. Multiply your rate by the typical response time in seconds and add a margin. If most responses take 3 seconds, 16 threads cap you near 5 per second; raise it to 32. - A different plan. Set
RATEto your plan's limit (50 on Starter, 100 on Growth) and raiseWORKERSto match. - Finding the URLs. The API fetches exactly the URL you send and does not follow links for you. To build the list, request a category page with
"links": trueand filter the result, as in our link crawler guide.
What it costs
A CSS extraction is 1 request unit per page on metered plans, and only delivered pages are billed; a 429 is never billed. On the Pool plan there is no request count at all, only the 10 per second limit this crawler is built around. See pricing.
Next steps
Read what the Pool plan is for to see whether unmetered fits your crawl. The error codes and rate limits are listed in the docs. Sign up free to get a key and run the crawler on a few pages.