Skip to content
All posts

How to turn a directory page into a lead list

Leads and local businesses5 min read
How to turn a directory page into a lead list: the post's first code sample

Industry directories, association member lists and "top 50 agencies" articles are some of the best lead sources on the web: someone has already done the work of picking the companies. This guide turns one of those pages into a spreadsheet of company websites with their public email addresses and phone numbers, in two steps.

The plan

  1. Fetch the directory page once with "links": true. You get every link on the page, already resolved to absolute URLs.
  2. Keep the links that point to company sites, then fetch each company's home page (and its contact page if the home page has no email) and read the mailto: and tel: links.

The API is stateless on purpose: it fetches the URL you give it and nothing else. It does not crawl a site for you. That keeps each step cheap and predictable, and you decide exactly which pages get fetched.

Step 1: get every link on the directory page

curl
curl https://scrape.land/v1/fetch \
  -H "X-Api-Key: YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com/top-50-agencies",
       "format": "text",
       "links": true}'

Asking for "format": "text" keeps the response small, since you only need the links. The response looks like this:

JSON
{
  "url": "https://example.com/top-50-agencies",
  "status": 200,
  "links": [
    "https://example.com/",
    "https://example.com/about",
    "https://agency-one.example.com/",
    "https://www.agency-two.example.com/",
    "https://example.com/top-50-agencies?page=2"
  ],
  "text": "Top 50 agencies\n1. Agency One ..."
}

Links are deduplicated, relative links are resolved (a <base href> on the page is honored, as a browser would), and mailto:, tel: and javascript: links are dropped from this list. Most directory pages link to their own navigation too, so the next step is to filter by domain.

Step 2: read contact details from each company

For each company link, use POST /v1/extract with two fields that collect every mailto: and tel: link on the page. The "all": true option returns every match as an array, and "attr": "href" reads the link target instead of its text. Add "links": true so you can find the contact page if the home page has no email.

curl
curl https://scrape.land/v1/extract \
  -H "X-Api-Key: YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://agency-one.example.com/",
       "links": true,
       "fields": {
         "name":   "title",
         "emails": {"css": "a[href^=\"mailto:\"]", "attr": "href", "all": true},
         "phones": {"css": "a[href^=\"tel:\"]", "attr": "href", "all": true}
       }}'
JSON
{
  "url": "https://agency-one.example.com/",
  "status": 200,
  "data": {
    "name": "Agency One | Brand and web design",
    "emails": ["mailto:hello@agency-one.example.com"],
    "phones": ["tel:+15125550142"]
  },
  "links": [
    "https://agency-one.example.com/work",
    "https://agency-one.example.com/contact"
  ]
}

The whole thing in Python

Python
import csv
from urllib.parse import urlparse
import requests

API = "https://scrape.land/v1"
HEAD = {"X-Api-Key": "YOUR_KEY"}
DIRECTORY = "https://example.com/top-50-agencies"
FIELDS = {
    "name": "title",
    "emails": {"css": "a[href^='mailto:']", "attr": "href", "all": True},
    "phones": {"css": "a[href^='tel:']", "attr": "href", "all": True},
}

def post(path, body):
    r = requests.post(API + path, headers=HEAD, json=body, timeout=90)
    r.raise_for_status()
    return r.json()

page = post("/fetch", {"url": DIRECTORY, "format": "text", "links": True})
own = urlparse(DIRECTORY).netloc
homes = {}
for link in page["links"]:
    host = urlparse(link).netloc
    if host and host != own and host not in homes:
        homes[host] = f"https://{host}/"

rows = []
for host, home in homes.items():
    try:
        res = post("/extract", {"url": home, "fields": FIELDS, "links": True})
    except requests.HTTPError as e:
        print("skip", host, e)
        continue
    data = res["data"]
    if not data["emails"]:
        contact = next((l for l in res.get("links", []) if "contact" in l.lower()), None)
        if contact:
            data = post("/extract", {"url": contact, "fields": FIELDS})["data"]
    rows.append({
        "company": data["name"] or host,
        "website": home,
        "email": ", ".join(sorted({e.removeprefix("mailto:").split("?")[0] for e in data["emails"] or []})),
        "phone": ", ".join(sorted({p.removeprefix("tel:") for p in data["phones"] or []})),
    })

with open("leads.csv", "w", newline="", encoding="utf-8") as f:
    w = csv.DictWriter(f, fieldnames=["company", "website", "email", "phone"])
    w.writeheader()
    w.writerows(rows)

Real directories need one more filter than this: drop links to social networks, app stores and the directory's own partners. A short blocklist of hosts is usually enough. If the directory is split over several pages, fetch each page URL you see in links yourself; there is no automatic pagination.

Cleaning the list

Before you import the CSV anywhere, remove duplicates by domain, since some directories list the same company twice under different names. Drop generic addresses you do not want to use, such as noreply@, and keep the website column: it is the easiest way to check a row by hand later.

When a site hides its email

Some sites show an email as plain text rather than a mailto: link, or only offer a contact form. For plain text, fetch the contact page with "format": "text" and match email addresses with a regular expression. If the contact page is built with JavaScript, add "render": true; the JavaScript guide explains the cost. A company with only a contact form has chosen not to publish an address, and it is best to respect that.

Only collect contact details that a business publishes for this purpose, and follow the email and privacy rules that apply to you and to the people you contact. Our Acceptable Use Policy covers what the service may be used for.

What it costs

Every fetch or CSS extraction is 1 request unit, and you only pay for responses that land (blocks and retries are free). A 50-company directory is about 51 to 101 units: one for the list, one per home page, and one more for each contact page you need. See pricing.

Next steps

The field syntax (attr, all, groups) is covered in the docs. If you do not have a directory to start from, find local businesses by category and city instead. Create a free account to get a key.

Start free with 1,000 requests Read the docs