How to scrape your first web page with curl
You do not need a scraping library, a headless browser or a list of IP addresses to get started. This guide takes you from no account to your first scraped page in about a minute: create a key, scrape your first web page with curl, read what comes back, and switch the output to Markdown or plain text with one parameter. Every response below is real output.
Step 1: get an API key
Create a free account. The Free plan includes 1,000 requests a month and needs no card. Confirm your email address from the link we send you: until you do, every request is refused with a 403 and the code email-unverified, so this step is not optional.
In the dashboard, open API keys and click + Create key. Copy the key somewhere safe. Every request sends it in the X-Api-Key header. An Authorization: Bearer header works too if your tooling prefers that.
To keep the key out of your shell history and out of the commands below, put it in an environment variable once:
export SCRAPELAND_KEY="YOUR_KEY"Step 2: scrape your first page with curl
The endpoint for "give me this page" is POST /v1/fetch. The body is JSON with one required field, url, which must be an absolute http or https URL. We fetch the page for you, rotating to a fresh exit IP and retrying past blocks, and hand the result back as JSON.
curl https://scrape.land/v1/fetch \
-H "X-Api-Key: $SCRAPELAND_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://example.com"}'This is the response we got, with the long HTML string shortened:
{
"html": "<!doctype html><html lang=en><head><meta charset=utf-8>…",
"status": 200,
"url": "https://example.com"
}What comes back
htmlis the page body, as the site sent it. The<and>sequences are just how JSON escapes<and>; any JSON parser turns them back into the real characters.statusis the status code the target site answered with. It is not the status of your API call. If you fetch a page that does not exist, the API call itself succeeds and you get"status": 404in the body, which is the site's real answer.urlis the URL that was fetched.
That split matters as soon as you write code around it. A non-200 HTTP status on the API call means our side refused or could not complete the request (a bad key, a bad URL, a rate limit). A non-200 status inside a 200 response means the site answered, and that is what it said.
To save just the HTML to a file, let jq unwrap it:
curl -s https://scrape.land/v1/fetch \
-H "X-Api-Key: $SCRAPELAND_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://example.com"}' | jq -r .html > example.htmlGet Markdown or plain text instead
Raw HTML is what you want if you are going to parse it yourself. Often you are not. Add format to the same request: html is the default, text returns the visible text, and markdown returns clean Markdown with the navigation, header, footer, scripts and styles stripped and links resolved to absolute URLs.
curl https://scrape.land/v1/fetch \
-H "X-Api-Key: $SCRAPELAND_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://example.com", "format": "markdown"}'{
"markdown": "This domain is for use in documentation examples without needing permission. This is not a service, avoid relying on it for testing and monitoring purposes.\n\n[Learn more](https://iana.org/help/example-domains)",
"status": 200,
"url": "https://example.com"
}The same page with "format": "text":
{
"status": 200,
"text": "Example Domain\nThis domain is for use in documentation examples without needing permission. This is not a service, avoid relying on it for testing and monitoring purposes.\nLearn more",
"url": "https://example.com"
}Notice that the result key follows the format: html, markdown or text. The text output also kept the page's <title> ("Example Domain") as its first line, while Markdown keeps only the body content.
Which one to use:
markdownfor anything going into a language model or a search index. Headings, lists, links and tables survive, and you do not pay model tokens for tags.textfor keyword checks, "does this page mention X" monitoring, or word counts.htmlwhen you want to run your own parser. If all you need is a few values, skip the parser and use/v1/extractwith selectors instead.
There is also raw, which returns the exact bytes base64-encoded with their content_type, for PDFs and images.
If your first request fails
These are the errors people hit on day one. The first two messages are exactly what the API returned when we triggered them:
401"invalid or missing API key": the header is missing, misspelled, or the variable is empty. Runecho $SCRAPELAND_KEYto check.400"url must be an absolute http or https URL": you sentexample.comwithouthttps://.403with"code": "email-unverified": confirm your email address first.- A shell error about quotes: on Windows
cmd.exe, single quotes do not work. Use Git Bash or WSL, or put the JSON in a file and send it with-d @body.json.
What it costs
One delivered page is 1 request unit, whatever the format: Markdown and text cost the same as HTML. Requests that fail or get blocked are not billed. The Free plan covers 1,000 a month. See pricing for paid plans.
Next steps
To pull specific values instead of the whole page, read how to extract fields with CSS selectors. For feeding pages to a language model, see clean Markdown for LLMs and RAG. Every option for this endpoint is in the fetch reference.