Skip to content
All posts

How to handle scraping errors and retries in Python

Getting started6 min read
How to handle scraping errors and retries in Python: the post's first code sample

A scraper that works on ten URLs usually breaks somewhere in the first ten thousand. The fix is not "retry everything three times": some errors go away on their own, some never will, and retrying the wrong ones wastes time or hides a bug. This guide shows how to handle scraping errors and retries in Python against the scrape.land API: what each status means, what you are billed for, and a small helper you can copy.

Two kinds of status

Every response has two layers, and mixing them up is the most common mistake.

Here is a real response for a page that does not exist. The API call returned HTTP 200; the site said 404:

JSON
{"data":{"t":"404 Not Found"},"status":404,"url":"https://books.toscrape.com/no-such-page.html"}

That is a correct answer, not a failure to retry. The page is gone. Log it and move on.

API status codes, and whether to retry

StatusMeaningRetry?
400The request itself is wrong: a URL without https://, an invalid XPath, a hostname that does not resolve.No. Fix the input.
401Missing or invalid API key.No.
402Quota or prepaid credit exhausted, or this key's own budget is used up.No. Top up, upgrade or raise the key budget.
403Destination not allowed, a feature your plan does not include, or an unconfirmed email address. The body says which.No.
413The response is larger than your plan's maximum response size.No.
429You went over your plan's rate limit. A Retry-After header says when to continue.Yes, after Retry-After.
502No working exit matched your filters, or the page could not be fetched.Yes, with backoff. Loosen country or max_latency if it persists.
503Momentarily at capacity, most often every browser slot busy on a render. Comes with Retry-After.Yes.
504The request ran past the 150-second budget for a synchronous call.Yes, once or twice. For slow renders, use an async job.

Refusals caused by your plan (402, the plan case of 403, 413, 429) all share one JSON shape with a stable code you can branch on, so you never have to parse the English sentence. This is what the Free plan returns when you ask for AI extraction:

JSON
{
  "capability": "AI extraction (prompt, schema, extract_type)",
  "code": "plan-upgrade-required",
  "error": "AI extraction (prompt, schema, extract_type) is not included on the \"free\" plan … Upgrade on the Plans page of your dashboard, or use CSS/XPath \"fields\" extraction, which every plan includes.",
  "fallback": "use CSS/XPath \"fields\" extraction, which every plan includes",
  "plan": "free",
  "required_plan": "scale"
}

The codes are plan-upgrade-required, quota-exhausted, key-budget-exhausted, rate-limited, response-too-large and email-unverified. Other errors carry just error.

Blocks: what we retry for you

When an exit IP gets a 403, a 429 or a 5xx from the target, we treat that as our exit failing, move to a different exit and try again, before you see anything. You do not need your own IP rotation. If you still see "status": 403 or 429 in a response body, the site is refusing this kind of request. One or two later retries with a pause can help, because a new request starts on fresh IPs; beyond that, look at country, render or a session rather than retrying harder.

What is and is not billed

So retrying a 404 costs you a unit each time for the same answer. Retrying a 429 or 503 costs nothing but time.

A Python helper

The helper below does four things: a connect and read timeout on every call, retries only on 429, 502, 503, 504 and network errors, exponential backoff that respects Retry-After, and an exception that carries the code for everything else.

Python
import logging
import random
import time

import requests

API = "https://scrape.land"
KEY = "YOUR_KEY"
RETRYABLE = {429, 502, 503, 504}   # worth another try; everything else is final

log = logging.getLogger("scrape")


class ScrapeError(Exception):
    def __init__(self, status, body):
        self.status = status
        self.code = body.get("code")
        super().__init__(f"{status} {self.code or 'error'}: {body.get('error')}")


def call(endpoint, payload, tries=4):
    for attempt in range(1, tries + 1):
        try:
            r = requests.post(API + endpoint, headers={"X-Api-Key": KEY},
                              json=payload, timeout=(10, 160))
        except (requests.ConnectionError, requests.Timeout) as e:
            if attempt == tries:
                raise
            wait = 2 ** attempt + random.random()
            log.warning("%s: network error %s, retry in %.1fs", payload.get("url"), e, wait)
            time.sleep(wait)
            continue

        if r.status_code == 200:
            return r.json()

        try:
            body = r.json()
        except ValueError:
            body = {"error": r.text[:200]}

        if r.status_code in RETRYABLE and attempt < tries:
            wait = float(r.headers.get("Retry-After") or 2 ** attempt) + random.random()
            log.warning("%s: %s %s, retry %d in %.1fs", payload.get("url"),
                        r.status_code, body.get("error"), attempt, wait)
            time.sleep(wait)
            continue

        raise ScrapeError(r.status_code, body)

The read timeout is 160 seconds because a synchronous call can legitimately take up to 150. The random fraction added to each wait (jitter) stops a hundred workers from retrying in the same instant.

Using it, and handling the site's own status separately:

Python
logging.basicConfig(level=logging.INFO, format="%(levelname)s %(message)s")

for url in ["https://books.toscrape.com/",
            "https://books.toscrape.com/no-such-page.html",
            "https://books-toscrape.invalid/"]:
    try:
        res = call("/v1/extract", {"url": url, "fields": {"title": "title"}})
        if res["status"] >= 400:
            log.info("%s: site answered %s", url, res["status"])
        else:
            log.info("%s: %s", url, res["data"])
    except ScrapeError as e:
        log.error("%s: %s", url, e)

Real output from our run:

shell
INFO https://books.toscrape.com/: {'title': 'All products | Books to Scrape - Sandbox'}
INFO https://books.toscrape.com/no-such-page.html: site answered 404
ERROR https://books-toscrape.invalid/: 400 error: could not resolve the target hostname … check the domain is spelled correctly and still exists

On an earlier run the second URL first timed out, and the helper did its job:

shell
WARNING https://books.toscrape.com/no-such-page.html: 504 the request took too long to complete; try a simpler page, fewer redirects, or retry, retry 1 in 2.2s
INFO https://books.toscrape.com/no-such-page.html: site answered 404

Logging that helps later

What it costs

Only delivered pages are billed: 1 request unit for a plain fetch or CSS/XPath extraction. Rate-limit refusals, busy responses and blocked attempts are free. See pricing.

Next steps

For long renders that run into the 150-second limit, read async jobs and webhooks. To cut the number of calls you make, see batch scraping many URLs in one call. The full error table is in the docs.

Start free with 1,000 requests Read the docs