Skip to content
All posts

How to use scrape.land with Scrapy

Integrations5 min read
How to use scrape.land with Scrapy: the post's first code sample

Scrapy already handles scheduling, retries and item pipelines. What it does not give you is a supply of exit IPs, or a way out of writing selectors for every site. There are two ways to use scrape.land with Scrapy: send the spider's own requests through the tunnel, or have the spider call the extraction API and get JSON back. This guide shows both, run for real, and the settings that keep either one inside your plan's rate limit.

Option 1: send the spider through the tunnel

Scrapy's built-in HttpProxyMiddleware reads meta["proxy"] on each request, including a username in the URL. Put your key there, with any options, and an empty password. Your parsing code does not change.

Python
import os
import scrapy

KEY = os.environ["SCRAPELAND_KEY"]
PROXY = f"http://{KEY}-country-us:@gateway.scrape.land:8080"


class QuotesSpider(scrapy.Spider):
    name = "quotes"
    custom_settings = {
        "CONCURRENT_REQUESTS": 8,
        "AUTOTHROTTLE_ENABLED": True,
        "AUTOTHROTTLE_TARGET_CONCURRENCY": 4,
        "DOWNLOAD_TIMEOUT": 60,
    }

    async def start(self):
        for n in (1, 2):
            url = f"https://quotes.toscrape.com/page/{n}/"
            yield scrapy.Request(url, meta={"proxy": PROXY})

    def parse(self, response):
        for q in response.css(".quote"):
            yield {
                "text": q.css(".text::text").get(),
                "author": q.css(".author::text").get(),
                "page": response.url,
            }
shell
export SCRAPELAND_KEY=YOUR_KEY
scrapy runspider quotes_proxy.py -O quotes.jsonl

In our run, the stats at the end showed 'downloader/response_status_count/200': 2 and 'item_scraped_count': 20, and the gateway's log showed both pages arriving through it. The first item:

JSON
{"text": "“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”", "author": "Albert Einstein", "page": "https://quotes.toscrape.com/page/1/"}

Two things to check in your own spider. First, async def start() is how Scrapy 2.13 and later begin a crawl; on older versions call it start_requests(). Recent Scrapy versions no longer call start_requests(), and a spider that also has start_urls then quietly fetches those without the tunnel. Second, the meta key must be on every request, including the ones you yield from parse when you follow links.

To set it once for the whole project, the scrapeland Python package (pip install scrapeland) ships a downloader middleware that does this for you, with per-request overrides:

Python
# settings.py
DOWNLOADER_MIDDLEWARES = {
    "scrapeland.scrapy.ScrapelandMiddleware": 740,
}
SCRAPELAND_API_KEY = "YOUR_KEY"
SCRAPELAND_COUNTRY = "us"            # optional default for every request

# per-request override:
yield scrapy.Request(url, meta={"scrapeland": {"country": "de", "session": "job42"}})

The options are the ones the tunnel documents: -country-, -session- (the same exit IP for about 10 minutes), -protocol- and -maxlatency-. Without a session, every request gets a fresh exit IP, which is usually what a crawl wants.

Option 2: call /v1/extract from the spider

Here Scrapy only talks to scrape.land. Each request is a POST with the target URL and a field map, and the response is JSON with the fields already extracted. This is the route to take when you want the API's features (Markdown, metadata, rendering on plans that include it) or simply want to stop maintaining XPath in Python.

Python
import os
import scrapy
from scrapy.http import JsonRequest

KEY = os.environ["SCRAPELAND_KEY"]
API = "https://scrape.land/v1/extract"
FIELDS = {"quotes": {"css": ".quote", "fields": {"text": ".text", "author": ".author"}}}


class QuotesApiSpider(scrapy.Spider):
    name = "quotes_api"
    custom_settings = {
        "CONCURRENT_REQUESTS": 16,
        "CONCURRENT_REQUESTS_PER_DOMAIN": 16,
        "DOWNLOAD_DELAY": 0.1,          # one slot (scrape.land), so ~10 requests/second
        "DOWNLOAD_TIMEOUT": 90,
        "RETRY_TIMES": 5,
    }

    async def start(self):
        for n in (1, 2):
            page = f"https://quotes.toscrape.com/page/{n}/"
            yield JsonRequest(API, data={"url": page, "fields": FIELDS},
                              headers={"X-Api-Key": KEY}, cb_kwargs={"page": page})

    def parse(self, response, page):
        body = response.json()
        for q in body["data"]["quotes"] or []:
            yield {**q, "page": page, "target_status": body["status"]}

This run also finished with two 200 responses and 20 items. JsonRequest sets the method, the JSON body and the Content-Type header. Pass the page URL along in cb_kwargs, because response.url is now the API endpoint, not the page. The body's status field is what the target site answered; keep it, so a 404 page is not mistaken for an empty one.

Settings that respect your plan's rate

Your plan's rate limit counts requests per second across your whole account: 10 on Pool, 50 on Starter, 100 on Growth. Go over it and you get a 429 with a Retry-After header, which is not billed. Scrapy's defaults retry 429 (along with 500, 502, 503 and 504), so a short overshoot costs time, not data. It is still better to pace the spider so you rarely hit it.

Scrapy's retries do not read Retry-After; with sensible pacing that rarely matters. If a tunnel refusal (such as 407 for a wrong key) arrives while opening an HTTPS connection, Scrapy reports it as a TunnelError and retries it like any connection error, so check the log if every request fails.

Which option to pick

Both use the same key and the same request allowance, so one project can mix them.

What it costs

A delivered request through the tunnel and a CSS extraction on the API are each 1 request unit; blocked and failed attempts are not billed. The Pool plan ($19 a month) has no request count, only its 10 per second limit, and includes both the tunnel and CSS extraction. See pricing.

Next steps

The tunnel options and the field map syntax are in the docs. For a crawler without Scrapy, see crawling thousands of pages at 10 requests a second. Sign up free to get a key and run either spider.

Start free with 1,000 requests Read the docs