Skip to content
All posts

How to extract every link on a page and build your own crawler

Leads and local businesses3 min read

Before you can scrape a site's pages, you need to know which pages exist. This tutorial gets every link on a page with one flag, already resolved to absolute URLs, and then builds a small same-site crawler in Python on top of it: breadth-first, deduplicated, with a hard page cap. The crawl logic stays in your code, so you decide exactly what gets fetched and what it costs.

Step 1: every link on one page

Add "links": true to a fetch. The response gets a links array: every <a href> on the page, resolved against the page URL (and its <base> tag, if it has one), deduplicated, and limited to http and https. mailto:, javascript: and in-page anchors are left out.

curl
curl https://scrape.land/v1/fetch \
  -H "X-Api-Key: YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://books.toscrape.com/", "format": "text", "links": true}'
Terminal showing a fetch of books.toscrape.com with links true, and a response listing absolute category URLs such as travel_2 and mystery_3, trimmed from 73 links in total
The real response from books.toscrape.com: 73 links, all absolute, even though the page's own hrefs are relative.

We asked for "format": "text" so the response carries the page's visible text instead of its full HTML, which keeps it small. Use "markdown" if you want to keep the page content for later, or "html" if you will run your own parser on it. A page returns at most 2,000 links.

Step 2: a crawler you control

The API is stateless: it reads the page you give it and never follows links on its own. That keeps crawling in your hands, where the decisions belong: which hosts are in scope, how deep to go, how many pages, and in what order. Here is a complete breadth-first crawler in about 30 lines:

Python
from collections import deque
from urllib.parse import urldefrag, urlparse

import requests

START = "https://quotes.toscrape.com/"
MAX_PAGES = 8  # every page is one request unit; keep a hard cap
host = urlparse(START).netloc

seen, queue, sitemap = {START}, deque([START]), []
while queue and len(sitemap) < MAX_PAGES:
    url = queue.popleft()
    r = requests.post("https://scrape.land/v1/fetch",
                      headers={"X-Api-Key": "YOUR_KEY"},
                      json={"url": url, "format": "text", "links": True},
                      timeout=60)
    if r.status_code != 200:
        print("skip", url, r.status_code)
        continue
    page = r.json()
    sitemap.append(url)
    for link in page.get("links", []):
        link = urldefrag(link).url
        if urlparse(link).netloc == host and link not in seen:
            seen.add(link)
            queue.append(link)

print(f"crawled {len(sitemap)} pages, {len(seen)} same-site URLs discovered")
for u in sitemap:
    print(" ", u)

We ran it against quotes.toscrape.com with an eight-page cap:

Terminal output of python crawl.py: crawled 8 pages, 49 same-site URLs discovered, followed by the list of crawled URLs such as the login page, the Albert Einstein author page and several tag pages
Eight fetches, 49 URLs discovered and queued for a later run.

Making the crawler useful

What it costs

"links": true adds nothing to a fetch: each page is 1 request unit. The eight-page crawl above cost 8 units. A page that needs JavaScript to build its links needs "render": true, which bills 5 units, or 1 with "block_resources": true. See pricing.

Next steps

The links option is documented in the docs. To turn a directory page into contact details, see directory page to lead list. For a site with numbered pages, the pagination tutorial is simpler than a full crawl. Create a free account and crawl your first site.

Start free with 1,000 requests Read the docs