How to extract every link on a page and build your own crawler
Before you can scrape a site's pages, you need to know which pages exist. This tutorial gets every link on a page with one flag, already resolved to absolute URLs, and then builds a small same-site crawler in Python on top of it: breadth-first, deduplicated, with a hard page cap. The crawl logic stays in your code, so you decide exactly what gets fetched and what it costs.
Step 1: every link on one page
Add "links": true to a fetch. The response gets a links array: every <a href> on the page, resolved against the page URL (and its <base> tag, if it has one), deduplicated, and limited to http and https. mailto:, javascript: and in-page anchors are left out.
curl https://scrape.land/v1/fetch \
-H "X-Api-Key: YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://books.toscrape.com/", "format": "text", "links": true}'
We asked for "format": "text" so the response carries the page's visible text instead of its full HTML, which keeps it small. Use "markdown" if you want to keep the page content for later, or "html" if you will run your own parser on it. A page returns at most 2,000 links.
Step 2: a crawler you control
The API is stateless: it reads the page you give it and never follows links on its own. That keeps crawling in your hands, where the decisions belong: which hosts are in scope, how deep to go, how many pages, and in what order. Here is a complete breadth-first crawler in about 30 lines:
from collections import deque
from urllib.parse import urldefrag, urlparse
import requests
START = "https://quotes.toscrape.com/"
MAX_PAGES = 8 # every page is one request unit; keep a hard cap
host = urlparse(START).netloc
seen, queue, sitemap = {START}, deque([START]), []
while queue and len(sitemap) < MAX_PAGES:
url = queue.popleft()
r = requests.post("https://scrape.land/v1/fetch",
headers={"X-Api-Key": "YOUR_KEY"},
json={"url": url, "format": "text", "links": True},
timeout=60)
if r.status_code != 200:
print("skip", url, r.status_code)
continue
page = r.json()
sitemap.append(url)
for link in page.get("links", []):
link = urldefrag(link).url
if urlparse(link).netloc == host and link not in seen:
seen.add(link)
queue.append(link)
print(f"crawled {len(sitemap)} pages, {len(seen)} same-site URLs discovered")
for u in sitemap:
print(" ", u)We ran it against quotes.toscrape.com with an eight-page cap:

Making the crawler useful
- Save the frontier. Write
queueandseento disk when the cap is reached, and the next run continues where this one stopped. - Filter before you fetch. A regex on the URL (
/product/,/catalogue/) keeps the crawl on the pages you want and skips login, cart and tag pages. This is where most of the cost saving is. - Extract while you crawl. The
linksarray comes back from/v1/fetchand/v1/batch. If you would rather crawl with/v1/extract, add a field such as"hrefs": {"css": "a", "attr": "href", "all": true}next to your data fields, and resolve the values withurljoin: one call then returns both the data and the next URLs. - Batch the fetches. Once you have a list of URLs, send up to 20 per call to
/v1/batchwith"links": true. See batch scraping. - Let a model pick. On a large hub page,
/v1/rankorders the links by relevance to a goal so you fetch the top few instead of all of them. See ranking links with AI. - Be polite. Respect
robots.txt, keep your concurrency modest, and do not crawl what you do not need.
What it costs
"links": true adds nothing to a fetch: each page is 1 request unit. The eight-page crawl above cost 8 units. A page that needs JavaScript to build its links needs "render": true, which bills 5 units, or 1 with "block_resources": true. See pricing.
Next steps
The links option is documented in the docs. To turn a directory page into contact details, see directory page to lead list. For a site with numbered pages, the pagination tutorial is simpler than a full crawl. Create a free account and crawl your first site.