How to use scrape.land with Scrapy
Scrapy already handles scheduling, retries and item pipelines. What it does not give you is a supply of exit IPs, or a way out of writing selectors for every site. There are two ways to use scrape.land with Scrapy: send the spider's own requests through the tunnel, or have the spider call the extraction API and get JSON back. This guide shows both, run for real, and the settings that keep either one inside your plan's rate limit.
Option 1: send the spider through the tunnel
Scrapy's built-in HttpProxyMiddleware reads meta["proxy"] on each request, including a username in the URL. Put your key there, with any options, and an empty password. Your parsing code does not change.
import os
import scrapy
KEY = os.environ["SCRAPELAND_KEY"]
PROXY = f"http://{KEY}-country-us:@gateway.scrape.land:8080"
class QuotesSpider(scrapy.Spider):
name = "quotes"
custom_settings = {
"CONCURRENT_REQUESTS": 8,
"AUTOTHROTTLE_ENABLED": True,
"AUTOTHROTTLE_TARGET_CONCURRENCY": 4,
"DOWNLOAD_TIMEOUT": 60,
}
async def start(self):
for n in (1, 2):
url = f"https://quotes.toscrape.com/page/{n}/"
yield scrapy.Request(url, meta={"proxy": PROXY})
def parse(self, response):
for q in response.css(".quote"):
yield {
"text": q.css(".text::text").get(),
"author": q.css(".author::text").get(),
"page": response.url,
}export SCRAPELAND_KEY=YOUR_KEY
scrapy runspider quotes_proxy.py -O quotes.jsonlIn our run, the stats at the end showed 'downloader/response_status_count/200': 2 and 'item_scraped_count': 20, and the gateway's log showed both pages arriving through it. The first item:
{"text": "“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”", "author": "Albert Einstein", "page": "https://quotes.toscrape.com/page/1/"}Two things to check in your own spider. First, async def start() is how Scrapy 2.13 and later begin a crawl; on older versions call it start_requests(). Recent Scrapy versions no longer call start_requests(), and a spider that also has start_urls then quietly fetches those without the tunnel. Second, the meta key must be on every request, including the ones you yield from parse when you follow links.
To set it once for the whole project, the scrapeland Python package (pip install scrapeland) ships a downloader middleware that does this for you, with per-request overrides:
# settings.py
DOWNLOADER_MIDDLEWARES = {
"scrapeland.scrapy.ScrapelandMiddleware": 740,
}
SCRAPELAND_API_KEY = "YOUR_KEY"
SCRAPELAND_COUNTRY = "us" # optional default for every request
# per-request override:
yield scrapy.Request(url, meta={"scrapeland": {"country": "de", "session": "job42"}})The options are the ones the tunnel documents: -country-, -session- (the same exit IP for about 10 minutes), -protocol- and -maxlatency-. Without a session, every request gets a fresh exit IP, which is usually what a crawl wants.
Option 2: call /v1/extract from the spider
Here Scrapy only talks to scrape.land. Each request is a POST with the target URL and a field map, and the response is JSON with the fields already extracted. This is the route to take when you want the API's features (Markdown, metadata, rendering on plans that include it) or simply want to stop maintaining XPath in Python.
import os
import scrapy
from scrapy.http import JsonRequest
KEY = os.environ["SCRAPELAND_KEY"]
API = "https://scrape.land/v1/extract"
FIELDS = {"quotes": {"css": ".quote", "fields": {"text": ".text", "author": ".author"}}}
class QuotesApiSpider(scrapy.Spider):
name = "quotes_api"
custom_settings = {
"CONCURRENT_REQUESTS": 16,
"CONCURRENT_REQUESTS_PER_DOMAIN": 16,
"DOWNLOAD_DELAY": 0.1, # one slot (scrape.land), so ~10 requests/second
"DOWNLOAD_TIMEOUT": 90,
"RETRY_TIMES": 5,
}
async def start(self):
for n in (1, 2):
page = f"https://quotes.toscrape.com/page/{n}/"
yield JsonRequest(API, data={"url": page, "fields": FIELDS},
headers={"X-Api-Key": KEY}, cb_kwargs={"page": page})
def parse(self, response, page):
body = response.json()
for q in body["data"]["quotes"] or []:
yield {**q, "page": page, "target_status": body["status"]}This run also finished with two 200 responses and 20 items. JsonRequest sets the method, the JSON body and the Content-Type header. Pass the page URL along in cb_kwargs, because response.url is now the API endpoint, not the page. The body's status field is what the target site answered; keep it, so a 404 page is not mistaken for an empty one.
Settings that respect your plan's rate
Your plan's rate limit counts requests per second across your whole account: 10 on Pool, 50 on Starter, 100 on Growth. Go over it and you get a 429 with a Retry-After header, which is not billed. Scrapy's defaults retry 429 (along with 500, 502, 503 and 504), so a short overshoot costs time, not data. It is still better to pace the spider so you rarely hit it.
- With the extraction API (option 2), every request goes to one host, so Scrapy puts them in one download slot and
DOWNLOAD_DELAYis effectively a global rate.0.1seconds is about 10 per second (Scrapy randomises each delay between half and one and a half times the setting, so it averages out). SetCONCURRENT_REQUESTS_PER_DOMAINhigh enough to keep that rate while each request waits for its page: rate times typical response time, plus a margin. - With the tunnel (option 1), Scrapy's slots are per target site, but the limit is per account. A spider crawling five sites with a delay of
0.1each can reach 50 per second. Cap the total withCONCURRENT_REQUESTS: requests per second is roughly concurrency divided by response time. - AutoThrottle adjusts the delay to how fast each site answers. It is kind to the sites you crawl, but it does not know your plan's limit, so keep the concurrency cap as well.
Scrapy's retries do not read Retry-After; with sensible pacing that rarely matters. If a tunnel refusal (such as 407 for a wrong key) arrives while opening an HTTPS connection, Scrapy reports it as a TunnelError and retries it like any connection error, so check the log if every request fails.
Which option to pick
- Tunnel when you already have a Scrapy project with working selectors and only need rotating or country-specific exits. Nothing else changes.
- Extraction API for new spiders, when you want JSON out without HTML parsing in Python, or when some pages need rendering.
Both use the same key and the same request allowance, so one project can mix them.
What it costs
A delivered request through the tunnel and a CSS extraction on the API are each 1 request unit; blocked and failed attempts are not billed. The Pool plan ($19 a month) has no request count, only its 10 per second limit, and includes both the tunnel and CSS extraction. See pricing.
Next steps
The tunnel options and the field map syntax are in the docs. For a crawler without Scrapy, see crawling thousands of pages at 10 requests a second. Sign up free to get a key and run either spider.