Skip to content
All posts

How to scrape websites with n8n, Make or Zapier

Integrations6 min read
How to scrape websites with n8n, Make or Zapier: the post's first code sample

Automation tools are good at moving data between apps and bad at reading web pages. The fix is one HTTP step: send a URL to the scrape.land extraction API and get back JSON with the fields you asked for, ready to map into a sheet, a CRM or a message. This guide shows the exact request to scrape websites with n8n, Make or Zapier, how to use the response, and how to handle errors.

The request, exactly

Every tool below ends up sending this. If you can make it work in curl, you can make it work anywhere:

curl
curl https://scrape.land/v1/extract \
  -H "X-Api-Key: YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://quotes.toscrape.com/",
       "fields": {
         "quotes": {"css": ".quote", "fields": {"text": ".text", "author": ".author"}}
       }}'

Each entry in fields is a name you choose and a CSS selector. quotes here is a group: .quote picks each quote box, and text and author are read inside each one. The real response, cut to two of its ten quotes:

JSON
{
  "data": {
    "quotes": [
      {"author": "Albert Einstein",
       "text": "“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”"},
      {"author": "J.K. Rowling",
       "text": "“It is our choices, Harry, that show what we truly are, far more than our abilities.”"}
    ]
  },
  "status": 200,
  "url": "https://quotes.toscrape.com/"
}

For one value per page (a product's name and price, an article's headline) use a flat map such as {"title": "h1", "price": ".price"} and you get data.title and data.price as plain strings. Flat maps are the easiest to use in all three tools.

n8n: the HTTP Request node

  1. Add an HTTP Request node. Set the method to POST and the URL to https://scrape.land/v1/extract.
  2. For authentication, create a generic header credential with the name X-Api-Key and your key as the value. Storing it as a credential keeps the key out of the workflow JSON when you share or export it.
  3. Turn on sending a body, choose JSON, and paste the body from above. To scrape a URL that an earlier node produced, replace the fixed URL with an expression such as {{ $json.url }}.
  4. Run the node. The output is the response JSON, so data.quotes is available to the next node. To turn the list into one item per quote, add a node that splits a field into items (Split Out in current versions) on data.quotes.

Make: the HTTP module

  1. Add the HTTP app's module for making a request. Set the URL and the method POST.
  2. Add two headers: X-Api-Key with your key and Content-Type with application/json.
  3. Choose a raw body with the JSON content type and paste the body. Map values from earlier modules into it with Make's mapping panel, for example the URL from a row of a Google Sheets module.
  4. Turn on parsing of the response, so later modules see data as fields instead of one long string. Run the module once so Make learns the structure.
  5. For a list such as data.quotes, add an Iterator on that array, and each quote becomes its own bundle for the modules after it.

Zapier: Webhooks by Zapier

  1. Add a Webhooks by Zapier action and pick the custom request event, which lets you send a raw body.
  2. Set the method to POST and the URL to https://scrape.land/v1/extract.
  3. Paste the JSON body into the data field. Insert fields from the trigger where the URL goes, keeping the quotes around it.
  4. Add the headers X-Api-Key and Content-Type: application/json.
  5. Test the step. The response fields appear for the next steps.

Zapier flattens nested lists in a response, so a group like quotes can arrive as comma-joined text. On Zapier, prefer a flat field map for single pages, or use a code step to loop over the list.

Scraping many URLs

Every tool can loop over rows and call the API once per URL. Two things help:

Handling errors

The API returns normal HTTP status codes with a JSON body that has an error message. Two real ones:

JSON
{"error": "invalid or missing API key"}
JSON
{"error": "url must be an absolute http or https URL"}

The first comes with 401, the second with 400 (a mapped field was empty, or not a full URL). By default all three tools stop the run on a status of 400 or above. Turn on the step's option to continue on error, or to return the response whatever the status, then branch on the status code:

A 200 from us does not mean the page was a success. Check the status field in the body: it is what the target site answered, so 404 there means the page is gone. A field whose selector matched nothing comes back as null; a filter step that skips rows with an empty key field keeps blanks out of your sheet.

Give the step a generous timeout (60 to 90 seconds). Most pages come back in a few seconds, but a slow site takes as long as it takes.

What it costs

A CSS extraction is 1 request unit per page, and only delivered pages are billed. The Free plan includes 1,000 requests a month with no card, enough to build and test a workflow. See pricing.

Next steps

Every parameter and response field is in the docs. If your destination is a spreadsheet, pulling scraped data into Google Sheets does it without an automation tool. For long jobs, async jobs and webhooks can call your workflow's webhook when a scrape finishes. Sign up free to get a key.

Start free with 1,000 requests Read the docs