How to find every page of a site from its sitemap
Before you scrape a site, you need its list of pages. Crawling link by link is slow and misses pages nothing links to. Most sites already publish that list for search engines: the sitemap. This guide shows how to find every page of a site from its sitemap: read robots.txt, fetch sitemap.xml and sitemap indexes through the API, parse the URLs in Python, then extract data from each page at a polite pace. The examples use djangoproject.com and all output is real.
Step 1: read robots.txt
robots.txt lives at the root of every site. It tells crawlers which paths to stay out of, sometimes how long to wait between requests (Crawl-delay), and often where the sitemap is (Sitemap: lines). Fetch it like any page:
curl https://scrape.land/v1/fetch \
-H "X-Api-Key: YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://www.djangoproject.com/robots.txt"}'{
"html": "User-agent: *\nDisallow: /admin\nDisallow: /checklists\n",
"status": 200,
"url": "https://www.djangoproject.com/robots.txt"
}The body comes back under html even though it is plain text: with the default format, the API returns the document exactly as the server sent it. This file has two Disallow rules and no Sitemap: line, so we fall back to the conventional location, /sitemap.xml. Many sites list one or more sitemaps in robots.txt, sometimes at a path you would never guess, so always check it first.
Step 2: fetch the sitemap
A sitemap is an XML file with one <url><loc> per page. Fetching it works the same way, and again the XML comes back untouched under html:
{
"html": "<?xml version=\"1.0\" encoding=\"UTF-8\"?>\n<urlset xmlns=\"http://www.sitemaps.org/schemas/sitemap/0.9\" …>\n<url><loc>https://www.djangoproject.com/weblog/2026/sep/24/dsf-member-of-the-month-ken-whitesell/</loc></url>…",
"status": 200,
"url": "https://www.djangoproject.com/sitemap.xml"
}The < sequences are JSON's escaping of <; your JSON parser undoes it.
Sitemap indexes
Large sites split their sitemap into several files and publish a sitemap index that lists them. It looks the same, except the root element is <sitemapindex> and each <loc> points to another sitemap instead of a page. The Django documentation site does this, with one sitemap per language.
If you only want the list of child sitemaps, you do not even need to parse XML yourself. /v1/extract reads the document with the same selector engine it uses for HTML, so loc works as a selector:
curl https://scrape.land/v1/extract \
-H "X-Api-Key: YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://docs.djangoproject.com/sitemap.xml",
"fields": {"sitemaps": {"css": "loc", "all": true}}}'{
"data": {
"sitemaps": [
"https://docs.djangoproject.com/sitemap-el.xml",
"https://docs.djangoproject.com/sitemap-en.xml",
"https://docs.djangoproject.com/sitemap-es.xml",
"https://docs.djangoproject.com/sitemap-fr.xml",
…
]
},
"status": 200,
"url": "https://docs.djangoproject.com/sitemap.xml"
}We cut the list after four entries. Each child is a normal sitemap you can fetch the same way. The XPath form, {"xpath": "//loc/text()", "all": true}, returns the same list.
Step 3: list every URL in Python
For a full run, parse the XML with the standard library. The script below reads robots.txt with urllib.robotparser (which also gives you the Sitemap: lines and any Crawl-delay), walks sitemap indexes up to two levels deep, keeps the URLs you care about, drops any that robots.txt disallows, and extracts the title and heading from each page.
import time
import xml.etree.ElementTree as ET
from urllib.parse import urlparse
from urllib.robotparser import RobotFileParser
import requests
API = "https://scrape.land"
KEY = "YOUR_KEY"
NS = {"sm": "http://www.sitemaps.org/schemas/sitemap/0.9"}
def fetch(url):
r = requests.post(f"{API}/v1/fetch", headers={"X-Api-Key": KEY},
json={"url": url}, timeout=60)
r.raise_for_status()
body = r.json()
return body["html"] if body["status"] == 200 else None
def read_robots(site):
rp = RobotFileParser()
rp.parse((fetch(site + "/robots.txt") or "").splitlines())
return rp
def page_urls(sitemap_url, depth=0):
xml = fetch(sitemap_url)
if not xml:
return []
root = ET.fromstring(xml.encode())
locs = [el.text.strip() for el in root.findall(".//sm:loc", NS)]
if root.tag.endswith("sitemapindex"):
if depth >= 2:
return []
urls = []
for child in locs:
urls += page_urls(child, depth + 1)
return urls
return locs
site = "https://www.djangoproject.com"
robots = read_robots(site)
sitemaps = robots.site_maps() or [site + "/sitemap.xml"]
delay = robots.crawl_delay("*") or 1
print("sitemaps:", sitemaps, "delay:", delay)
urls = []
for sm in sitemaps:
urls += page_urls(sm)
print(len(urls), "URLs in the sitemap")
wanted = [u for u in urls
if urlparse(u).path.startswith("/foundation/") and robots.can_fetch("*", u)]
print(len(wanted), "under /foundation/")
for u in wanted[:3]:
r = requests.post(f"{API}/v1/extract", headers={"X-Api-Key": KEY},
json={"url": u, "fields": {"title": "title", "h1": "h1"}},
timeout=60)
print(r.json())
time.sleep(delay)Real output, with the three extraction results shown as the script printed them:
sitemaps: ['https://www.djangoproject.com/sitemap.xml'] delay: 1
1017 URLs in the sitemap
91 under /foundation/
{'data': {'h1': 'About the Django Software Foundation', 'title': 'About the Django Software Foundation | Django'}, 'status': 200, 'url': 'https://www.djangoproject.com/foundation/'}
{'data': {'h1': 'Contributor License Agreements', 'title': 'Contributor License Agreements | Django'}, 'status': 200, 'url': 'https://www.djangoproject.com/foundation/cla/'}
{'data': {'h1': 'About the Django Contributor License Agreement', 'title': 'About the Django Contributor License Agreement | Django'}, 'status': 200, 'url': 'https://www.djangoproject.com/foundation/cla/faq/'}Two requests (robots.txt and the sitemap) produced a list of 1,017 URLs. Filtering before you extract anything is where the savings are: this sitemap is mostly blog posts, and we only wanted the 91 foundation pages. Filter on urlparse(u).path rather than a substring of the whole URL; a plain "/foundation/" in u also matched a 2008 blog post with "foundation" in its address.
A few details the script handles, and a few it leaves to you:
- Namespaces. Sitemap tags live in the
http://www.sitemaps.org/schemas/sitemap/0.9namespace, so a barefindall("loc")finds nothing. TheNSmap fixes that. - Depth limit. Indexes should only point to sitemaps, but the cap of two levels means a misconfigured site cannot send the script in circles.
- Gzipped sitemaps. Some sites publish
sitemap.xml.gz. Fetch those with"format": "raw", which returns the exact bytes base64-encoded underbase64, then run them throughbase64.b64decodeandgzip.decompress. <lastmod>. Many sitemaps include a last-modified date per URL. On repeat runs, compare it with your previous run and only re-scrape pages that changed.
Step 4: extract politely
A sitemap hands you every URL at once, which makes it easy to hit a site harder than you mean to. Rotating IPs spreads your requests across exits, but the same site still serves every one of them. Some habits that keep you welcome:
- Respect
Disallow. The script checksrobots.can_fetchfor every URL before requesting it. - Pace yourself. Use the site's
Crawl-delaywhen it gives one, and a sensible default (the script uses one second) when it does not. A small site does not need a hundred parallel workers. - Only fetch what you need. Filter the URL list first. Every page you skip is one less request for them and one less unit for you.
- Batch when you do go faster.
POST /v1/batchtakes up to 20 URLs and the samefieldsin one call, which is simpler than managing your own thread pool.
If a site has no sitemap at all (neither in robots.txt nor at /sitemap.xml), fall back to discovering pages by links: add "links": true to a fetch and you get every link on the page as an absolute URL.
What it costs
Each robots.txt, sitemap or page fetch is 1 request unit, and so is each CSS or XPath extraction. For this example: 2 units to list 1,017 URLs, then 1 per page you extract. Failed and blocked requests are not billed. See pricing.
Next steps
For sites without a sitemap, read how to extract every link and build a crawler. To extract from many URLs in fewer calls, see batch scraping. Every fetch option is in the docs.