How to handle scraping errors and retries in Python
A scraper that works on ten URLs usually breaks somewhere in the first ten thousand. The fix is not "retry everything three times": some errors go away on their own, some never will, and retrying the wrong ones wastes time or hides a bug. This guide shows how to handle scraping errors and retries in Python against the scrape.land API: what each status means, what you are billed for, and a small helper you can copy.
Two kinds of status
Every response has two layers, and mixing them up is the most common mistake.
- The HTTP status of your API call.
200means we fetched the page and are handing you the result. Anything else means we refused the request or could not complete it, and the JSON body has anerrormessage saying why. - The
statusfield inside a 200 response. That is what the target site answered. It is passed through unchanged.
Here is a real response for a page that does not exist. The API call returned HTTP 200; the site said 404:
{"data":{"t":"404 Not Found"},"status":404,"url":"https://books.toscrape.com/no-such-page.html"}That is a correct answer, not a failure to retry. The page is gone. Log it and move on.
API status codes, and whether to retry
| Status | Meaning | Retry? |
|---|---|---|
400 | The request itself is wrong: a URL without https://, an invalid XPath, a hostname that does not resolve. | No. Fix the input. |
401 | Missing or invalid API key. | No. |
402 | Quota or prepaid credit exhausted, or this key's own budget is used up. | No. Top up, upgrade or raise the key budget. |
403 | Destination not allowed, a feature your plan does not include, or an unconfirmed email address. The body says which. | No. |
413 | The response is larger than your plan's maximum response size. | No. |
429 | You went over your plan's rate limit. A Retry-After header says when to continue. | Yes, after Retry-After. |
502 | No working exit matched your filters, or the page could not be fetched. | Yes, with backoff. Loosen country or max_latency if it persists. |
503 | Momentarily at capacity, most often every browser slot busy on a render. Comes with Retry-After. | Yes. |
504 | The request ran past the 150-second budget for a synchronous call. | Yes, once or twice. For slow renders, use an async job. |
Refusals caused by your plan (402, the plan case of 403, 413, 429) all share one JSON shape with a stable code you can branch on, so you never have to parse the English sentence. This is what the Free plan returns when you ask for AI extraction:
{
"capability": "AI extraction (prompt, schema, extract_type)",
"code": "plan-upgrade-required",
"error": "AI extraction (prompt, schema, extract_type) is not included on the \"free\" plan … Upgrade on the Plans page of your dashboard, or use CSS/XPath \"fields\" extraction, which every plan includes.",
"fallback": "use CSS/XPath \"fields\" extraction, which every plan includes",
"plan": "free",
"required_plan": "scale"
}The codes are plan-upgrade-required, quota-exhausted, key-budget-exhausted, rate-limited, response-too-large and email-unverified. Other errors carry just error.
Blocks: what we retry for you
When an exit IP gets a 403, a 429 or a 5xx from the target, we treat that as our exit failing, move to a different exit and try again, before you see anything. You do not need your own IP rotation. If you still see "status": 403 or 429 in a response body, the site is refusing this kind of request. One or two later retries with a pause can help, because a new request starts on fresh IPs; beyond that, look at country, render or a session rather than retrying harder.
What is and is not billed
- Not billed: blocked attempts, requests where every exit failed,
429over your rate limit,503busy responses, responses refused with413, and a render or AI extraction that fails. - Billed: a delivered page, including a genuine error page from the site such as a
404. It is the site's real answer and it took a request to get it.
So retrying a 404 costs you a unit each time for the same answer. Retrying a 429 or 503 costs nothing but time.
A Python helper
The helper below does four things: a connect and read timeout on every call, retries only on 429, 502, 503, 504 and network errors, exponential backoff that respects Retry-After, and an exception that carries the code for everything else.
import logging
import random
import time
import requests
API = "https://scrape.land"
KEY = "YOUR_KEY"
RETRYABLE = {429, 502, 503, 504} # worth another try; everything else is final
log = logging.getLogger("scrape")
class ScrapeError(Exception):
def __init__(self, status, body):
self.status = status
self.code = body.get("code")
super().__init__(f"{status} {self.code or 'error'}: {body.get('error')}")
def call(endpoint, payload, tries=4):
for attempt in range(1, tries + 1):
try:
r = requests.post(API + endpoint, headers={"X-Api-Key": KEY},
json=payload, timeout=(10, 160))
except (requests.ConnectionError, requests.Timeout) as e:
if attempt == tries:
raise
wait = 2 ** attempt + random.random()
log.warning("%s: network error %s, retry in %.1fs", payload.get("url"), e, wait)
time.sleep(wait)
continue
if r.status_code == 200:
return r.json()
try:
body = r.json()
except ValueError:
body = {"error": r.text[:200]}
if r.status_code in RETRYABLE and attempt < tries:
wait = float(r.headers.get("Retry-After") or 2 ** attempt) + random.random()
log.warning("%s: %s %s, retry %d in %.1fs", payload.get("url"),
r.status_code, body.get("error"), attempt, wait)
time.sleep(wait)
continue
raise ScrapeError(r.status_code, body)The read timeout is 160 seconds because a synchronous call can legitimately take up to 150. The random fraction added to each wait (jitter) stops a hundred workers from retrying in the same instant.
Using it, and handling the site's own status separately:
logging.basicConfig(level=logging.INFO, format="%(levelname)s %(message)s")
for url in ["https://books.toscrape.com/",
"https://books.toscrape.com/no-such-page.html",
"https://books-toscrape.invalid/"]:
try:
res = call("/v1/extract", {"url": url, "fields": {"title": "title"}})
if res["status"] >= 400:
log.info("%s: site answered %s", url, res["status"])
else:
log.info("%s: %s", url, res["data"])
except ScrapeError as e:
log.error("%s: %s", url, e)Real output from our run:
INFO https://books.toscrape.com/: {'title': 'All products | Books to Scrape - Sandbox'}
INFO https://books.toscrape.com/no-such-page.html: site answered 404
ERROR https://books-toscrape.invalid/: 400 error: could not resolve the target hostname … check the domain is spelled correctly and still existsOn an earlier run the second URL first timed out, and the helper did its job:
WARNING https://books.toscrape.com/no-such-page.html: 504 the request took too long to complete; try a simpler page, fewer redirects, or retry, retry 1 in 2.2s
INFO https://books.toscrape.com/no-such-page.html: site answered 404Logging that helps later
- Log the URL, the API status and the
codeon every failure. A spike ofrate-limitedmeans slow down; a spike of 404s means your URL list is stale. - Keep the
X-Request-Idresponse header for failures. It identifies the request if you need to ask support about it. - Stop the whole run on
401and402. Every later request will fail the same way. CheckGET /v1/accountbefore a large job to see how many included requests you have left. - Watch
field_errorson extraction responses. A 200 with a broken selector is still a problem, just a quieter one.
What it costs
Only delivered pages are billed: 1 request unit for a plain fetch or CSS/XPath extraction. Rate-limit refusals, busy responses and blocked attempts are free. See pricing.
Next steps
For long renders that run into the 150-second limit, read async jobs and webhooks. To cut the number of calls you make, see batch scraping many URLs in one call. The full error table is in the docs.