Skip to content
All posts

How to get clean Markdown from any page for an LLM

Getting started4 min read

Language models read Markdown well and raw HTML badly. HTML is full of navigation, scripts, class names and tracking tags that cost tokens and add noise. This guide fetches any web page as clean Markdown, ready to paste into a prompt or split into chunks for a retrieval (RAG) index.

The request

Call POST /v1/fetch with "format": "markdown":

curl
curl https://scrape.land/v1/fetch \
  -H "X-Api-Key: YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://news.example.com/2026/09/city-opens-new-library",
       "format": "markdown",
       "metadata": true}'
JSON
{
  "url": "https://news.example.com/2026/09/city-opens-new-library",
  "status": 200,
  "metadata": {
    "title": "City opens new library",
    "description": "The central library reopens after a two-year renovation.",
    "canonical": "https://news.example.com/2026/09/city-opens-new-library"
  },
  "markdown": "# City opens new library\n\nThe central library reopened on Monday after a two-year renovation.\n\n## What changed\n\n- A new children's floor\n- Longer opening hours\n\n![Reading room](https://news.example.com/img/reading-room.jpg)\n\nRead the [council announcement](https://news.example.com/council/library)."
}

What the conversion does:

This is what the conversion looks like on a real page. We fetched a product page from books.toscrape.com, a practice shop, as Markdown:

Terminal showing a fetch with format markdown for the A Light in the Attic product page, and the Markdown output: breadcrumb links, the cover image, a heading, price, stock, the product description and a Markdown table with UPC, prices, tax and availability
The real output (the long description trimmed). Breadcrumb and image links are absolute, and the product information table survives as a Markdown table.

"metadata": true is optional. It adds the page title, description, canonical URL, OpenGraph tags and schema.org JSON-LD, which make good metadata for each chunk in your index (source URL, title, publish date).

From Markdown to chunks in Python

A simple and effective way to chunk is by heading: each section becomes a chunk, prefixed with the page title so it still makes sense on its own.

Python
import json
import re
import requests

def page_markdown(url):
    r = requests.post(
        "https://scrape.land/v1/fetch",
        headers={"X-Api-Key": "YOUR_KEY"},
        json={"url": url, "format": "markdown", "metadata": True},
        timeout=60,
    )
    r.raise_for_status()
    return r.json()

def chunks(doc, max_chars=2000):
    title = (doc.get("metadata") or {}).get("title") or doc["url"]
    sections = re.split(r"\n(?=#{1,3} )", doc["markdown"])
    for section in sections:
        section = section.strip()
        for i in range(0, len(section), max_chars):
            yield {
                "source": doc["url"],
                "title": title,
                "text": f"{title}\n\n{section[i:i + max_chars]}",
            }

urls = [
    "https://news.example.com/2026/09/city-opens-new-library",
    "https://example.com/docs/getting-started",
]
with open("chunks.jsonl", "w", encoding="utf-8") as f:
    for url in urls:
        for c in chunks(page_markdown(url)):
            f.write(json.dumps(c, ensure_ascii=False) + "\n")

chunks.jsonl is one chunk per line, ready for whatever embedding model and vector store you use. Splitting by characters is crude; if your pages have very long sections, split on paragraphs (\n\n) inside each section instead.

Pages that need a browser

If the Markdown comes back nearly empty, the page probably builds its content with JavaScript. Add "render": true and "block_resources": true to the same request. The JavaScript guide covers when that is needed and what it costs.

Markdown, text or HTML?

If you want specific facts from the page rather than all of its text, skip the chunking and ask for them directly with CSS fields or an AI prompt.

Common problems

A note on scale

For many URLs, send them in parallel from your own code or use POST /v1/batch, which takes several URLs per call. The API does not crawl: you give it the list of URLs, it turns each one into Markdown. That keeps you in control of what goes into your corpus. Respect each site's terms and our Acceptable Use Policy.

What it costs

Markdown is the same price as any fetch: 1 request unit per page, on every plan including Free (1,000 requests a month, no card). There is no extra charge for the conversion. Rendering, if you need it, is 5 units or 1 with block_resources. See pricing.

Next steps

All /v1/fetch formats and options are in the docs. To let an assistant fetch pages as Markdown on its own, see using scrape.land from Claude over MCP. Sign up free to get a key.

Start free with 1,000 requests Read the docs