Skip to content
All posts

How to find every page of a site from its sitemap

Getting started6 min read
How to find every page of a site from its sitemap: the post's first code sample

Before you scrape a site, you need its list of pages. Crawling link by link is slow and misses pages nothing links to. Most sites already publish that list for search engines: the sitemap. This guide shows how to find every page of a site from its sitemap: read robots.txt, fetch sitemap.xml and sitemap indexes through the API, parse the URLs in Python, then extract data from each page at a polite pace. The examples use djangoproject.com and all output is real.

Step 1: read robots.txt

robots.txt lives at the root of every site. It tells crawlers which paths to stay out of, sometimes how long to wait between requests (Crawl-delay), and often where the sitemap is (Sitemap: lines). Fetch it like any page:

curl
curl https://scrape.land/v1/fetch \
  -H "X-Api-Key: YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://www.djangoproject.com/robots.txt"}'
JSON
{
  "html": "User-agent: *\nDisallow: /admin\nDisallow: /checklists\n",
  "status": 200,
  "url": "https://www.djangoproject.com/robots.txt"
}

The body comes back under html even though it is plain text: with the default format, the API returns the document exactly as the server sent it. This file has two Disallow rules and no Sitemap: line, so we fall back to the conventional location, /sitemap.xml. Many sites list one or more sitemaps in robots.txt, sometimes at a path you would never guess, so always check it first.

Step 2: fetch the sitemap

A sitemap is an XML file with one <url><loc> per page. Fetching it works the same way, and again the XML comes back untouched under html:

JSON
{
  "html": "<?xml version=\"1.0\" encoding=\"UTF-8\"?>\n<urlset xmlns=\"http://www.sitemaps.org/schemas/sitemap/0.9\" …>\n<url><loc>https://www.djangoproject.com/weblog/2026/sep/24/dsf-member-of-the-month-ken-whitesell/</loc></url>…",
  "status": 200,
  "url": "https://www.djangoproject.com/sitemap.xml"
}

The < sequences are JSON's escaping of <; your JSON parser undoes it.

Sitemap indexes

Large sites split their sitemap into several files and publish a sitemap index that lists them. It looks the same, except the root element is <sitemapindex> and each <loc> points to another sitemap instead of a page. The Django documentation site does this, with one sitemap per language.

If you only want the list of child sitemaps, you do not even need to parse XML yourself. /v1/extract reads the document with the same selector engine it uses for HTML, so loc works as a selector:

curl
curl https://scrape.land/v1/extract \
  -H "X-Api-Key: YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://docs.djangoproject.com/sitemap.xml",
       "fields": {"sitemaps": {"css": "loc", "all": true}}}'
JSON
{
  "data": {
    "sitemaps": [
      "https://docs.djangoproject.com/sitemap-el.xml",
      "https://docs.djangoproject.com/sitemap-en.xml",
      "https://docs.djangoproject.com/sitemap-es.xml",
      "https://docs.djangoproject.com/sitemap-fr.xml",
      …
    ]
  },
  "status": 200,
  "url": "https://docs.djangoproject.com/sitemap.xml"
}

We cut the list after four entries. Each child is a normal sitemap you can fetch the same way. The XPath form, {"xpath": "//loc/text()", "all": true}, returns the same list.

Step 3: list every URL in Python

For a full run, parse the XML with the standard library. The script below reads robots.txt with urllib.robotparser (which also gives you the Sitemap: lines and any Crawl-delay), walks sitemap indexes up to two levels deep, keeps the URLs you care about, drops any that robots.txt disallows, and extracts the title and heading from each page.

Python
import time
import xml.etree.ElementTree as ET
from urllib.parse import urlparse
from urllib.robotparser import RobotFileParser

import requests

API = "https://scrape.land"
KEY = "YOUR_KEY"
NS = {"sm": "http://www.sitemaps.org/schemas/sitemap/0.9"}


def fetch(url):
    r = requests.post(f"{API}/v1/fetch", headers={"X-Api-Key": KEY},
                      json={"url": url}, timeout=60)
    r.raise_for_status()
    body = r.json()
    return body["html"] if body["status"] == 200 else None


def read_robots(site):
    rp = RobotFileParser()
    rp.parse((fetch(site + "/robots.txt") or "").splitlines())
    return rp


def page_urls(sitemap_url, depth=0):
    xml = fetch(sitemap_url)
    if not xml:
        return []
    root = ET.fromstring(xml.encode())
    locs = [el.text.strip() for el in root.findall(".//sm:loc", NS)]
    if root.tag.endswith("sitemapindex"):
        if depth >= 2:
            return []
        urls = []
        for child in locs:
            urls += page_urls(child, depth + 1)
        return urls
    return locs


site = "https://www.djangoproject.com"
robots = read_robots(site)
sitemaps = robots.site_maps() or [site + "/sitemap.xml"]
delay = robots.crawl_delay("*") or 1
print("sitemaps:", sitemaps, "delay:", delay)

urls = []
for sm in sitemaps:
    urls += page_urls(sm)
print(len(urls), "URLs in the sitemap")

wanted = [u for u in urls
          if urlparse(u).path.startswith("/foundation/") and robots.can_fetch("*", u)]
print(len(wanted), "under /foundation/")

for u in wanted[:3]:
    r = requests.post(f"{API}/v1/extract", headers={"X-Api-Key": KEY},
                      json={"url": u, "fields": {"title": "title", "h1": "h1"}},
                      timeout=60)
    print(r.json())
    time.sleep(delay)

Real output, with the three extraction results shown as the script printed them:

shell
sitemaps: ['https://www.djangoproject.com/sitemap.xml'] delay: 1
1017 URLs in the sitemap
91 under /foundation/
{'data': {'h1': 'About the Django Software Foundation', 'title': 'About the Django Software Foundation | Django'}, 'status': 200, 'url': 'https://www.djangoproject.com/foundation/'}
{'data': {'h1': 'Contributor License Agreements', 'title': 'Contributor License Agreements | Django'}, 'status': 200, 'url': 'https://www.djangoproject.com/foundation/cla/'}
{'data': {'h1': 'About the Django Contributor License Agreement', 'title': 'About the Django Contributor License Agreement | Django'}, 'status': 200, 'url': 'https://www.djangoproject.com/foundation/cla/faq/'}

Two requests (robots.txt and the sitemap) produced a list of 1,017 URLs. Filtering before you extract anything is where the savings are: this sitemap is mostly blog posts, and we only wanted the 91 foundation pages. Filter on urlparse(u).path rather than a substring of the whole URL; a plain "/foundation/" in u also matched a 2008 blog post with "foundation" in its address.

A few details the script handles, and a few it leaves to you:

Step 4: extract politely

A sitemap hands you every URL at once, which makes it easy to hit a site harder than you mean to. Rotating IPs spreads your requests across exits, but the same site still serves every one of them. Some habits that keep you welcome:

If a site has no sitemap at all (neither in robots.txt nor at /sitemap.xml), fall back to discovering pages by links: add "links": true to a fetch and you get every link on the page as an absolute URL.

What it costs

Each robots.txt, sitemap or page fetch is 1 request unit, and so is each CSS or XPath extraction. For this example: 2 units to list 1,017 URLs, then 1 per page you extract. Failed and blocked requests are not billed. See pricing.

Next steps

For sites without a sitemap, read how to extract every link and build a crawler. To extract from many URLs in fewer calls, see batch scraping. Every fetch option is in the docs.

Start free with 1,000 requests Read the docs