Skip to content
All posts

How to collect contact details from websites in Python

Leads and local businesses6 min read
How to collect contact details from websites in Python: the post's first code sample

You have a column of company websites, perhaps from a map search or a trade directory, and you need an email and a phone number for each. This guide builds a small Python script that collects contact details from a list of websites: it fetches each homepage through POST /v1/fetch, falls back to the contact page when it has to, and writes a CSV. It is the same approach the dashboard's Lead Finder uses.

The plan

For every site:

  1. Fetch the homepage as HTML with "links": true, which adds a links array of every link on the page as an absolute URL, and "metadata": true, which adds a metadata object with the page title (used as the company name).
  2. Parse the HTML for mailto: links, plain addresses in the text, Cloudflare-protected addresses, tel: links, the schema.org telephone field and social profile links.
  3. If an email or a phone is still missing, pick one contact or about page on the same host from links, fetch it and parse it the same way.

That is one or two fetches per company. Each delivered fetch is 1 request unit; a fetch that fails is not billed.

What a fetch returns

A request with both options looks like this:

curl
curl https://scrape.land/v1/fetch \
  -H "X-Api-Key: YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://www.greenhalghs.com/",
       "format": "html", "metadata": true, "links": true}'

The response has url, status and html, plus metadata and links. For this bakery's homepage the title came back as "Greenhalgh's Craft Bakery – Fresh Breads, Pies & Cakes", and among the links were https://www.greenhalghs.com/contact-us/ and https://www.greenhalghs.com/about-us-greenhalghs-craft-bakery/. The script prefers "contact" over "about" when both exist, because the about page rarely carries an email.

The script

Python 3.9 or newer, plus requests. It reads four sites at a time with a thread pool.

Python
import csv, re
from concurrent.futures import ThreadPoolExecutor
from urllib.parse import urlparse, unquote
import requests

API = "https://scrape.land/v1/fetch"
KEY = "YOUR_KEY"
SITES = ["https://oldroydgroup.co.uk/", "https://wbtrade.co.uk/",
         "https://www.greenhalghs.com/", "https://josiescakes.co.uk/"]

EMAIL = re.compile(r"[A-Z0-9._%+-]+@[A-Z0-9.-]+\.[A-Z]{2,}", re.I)
WEBMAIL = re.compile(r"@(gmail|googlemail|yahoo|hotmail|outlook|live|aol|icloud)\.", re.I)
JUNK = re.compile(r"\.(png|jpe?g|gif|svg|webp|css|js)$|@(example|domain|sentry|wixpress)\.", re.I)
CF = re.compile(r'(?:data-cfemail=["\']|/cdn-cgi/l/email-protection#)([0-9a-f]+)', re.I)
TEL = re.compile(r'href=["\']tel:([^"\']+)["\']', re.I)
LD_TEL = re.compile(r'"telephone"\s*:\s*"([^"]+)"', re.I)
SOCIAL = re.compile(r'href=["\'](https?://(?:www\.)?(?:linkedin\.com/company|facebook\.com'
                    r'|instagram\.com|x\.com|twitter\.com)/[^"\'?#\s]+)', re.I)

def fetch(url, **opts):
    r = requests.post(API, headers={"X-Api-Key": KEY},
                      json={"url": url, "format": "html", **opts}, timeout=90)
    r.raise_for_status()
    return r.json()

def cf_decode(hexstr):
    # Cloudflare's email obfuscation: first byte is the XOR key.
    key = int(hexstr[:2], 16)
    return "".join(chr(int(hexstr[i:i + 2], 16) ^ key) for i in range(2, len(hexstr), 2))

def parse(html):
    html = html.replace("@", "@")
    emails = {unquote(m) for m in re.findall(r"mailto:([^\"'?>\s]+)", html, re.I)}
    emails |= {cf_decode(h) for h in CF.findall(html)}
    visible = re.sub(r"<script(?![^>]*ld\+json).*?</script>|<style.*?</style>", " ",
                     html, flags=re.S | re.I)
    emails |= set(EMAIL.findall(visible))
    emails = {e.lower() for e in emails
              if EMAIL.fullmatch(e) and not WEBMAIL.search(e) and not JUNK.search(e)}
    phones = {re.sub(r"\D", "", p): unquote(p).strip()      # one entry per number
              for p in TEL.findall(html) + LD_TEL.findall(html)}
    return emails, set(phones.values()), set(SOCIAL.findall(html))

def contact_page(links, host):
    same = [l for l in links if (urlparse(l).hostname or "").removeprefix("www.") == host]
    for word in ("contact", "kontakt", "about", "impressum"):
        for l in same:
            if word in urlparse(l).path.lower():
                return l
    return None

def lead(site):
    host = urlparse(site).hostname.removeprefix("www.")
    row = {"website": site, "company": host, "emails": set(), "phones": set(),
           "social": set(), "units": 0, "error": ""}
    try:
        d = fetch(site, metadata=True, links=True)
        row["units"] += 1
        title = (d.get("metadata") or {}).get("title") or ""
        row["company"] = " ".join(title.split()) or host
        row["emails"], row["phones"], row["social"] = parse(d.get("html", ""))
        page = contact_page(d.get("links", []), host)
        if page and (not row["emails"] or not row["phones"]):
            e, p, s = parse(fetch(page).get("html", ""))
            row["units"] += 1
            row["emails"] |= e
            row["phones"] |= p
            row["social"] |= s
    except requests.HTTPError as e:
        row["error"] = f'{e.response.status_code} {e.response.json().get("error", "")}'
    return row

with ThreadPoolExecutor(max_workers=4) as pool:
    rows = list(pool.map(lead, SITES))

with open("contacts.csv", "w", newline="", encoding="utf-8") as f:
    w = csv.writer(f)
    w.writerow(["company", "website", "emails", "phones", "social", "error"])
    for r in rows:
        w.writerow([r["company"], r["website"], "; ".join(sorted(r["emails"])),
                    "; ".join(sorted(r["phones"])), "; ".join(sorted(r["social"])), r["error"]])
        print(f'{r["website"]:32} emails={sorted(r["emails"])} phones={sorted(r["phones"])} '
              f'social={len(r["social"])} {r["error"]}')

print(sum(r["units"] for r in rows), "request units used")

Some notes on the parsing:

A real run

We ran the script against four small UK business websites taken from a map search. It printed:

shell
https://oldroydgroup.co.uk/      emails=[] phones=[] social=0
https://wbtrade.co.uk/           emails=['sales@wbtrade.co.uk'] phones=['01924267488'] social=0
https://www.greenhalghs.com/     emails=[] phones=['01204 696204'] social=3
https://josiescakes.co.uk/       emails=[] phones=[] social=0
5 request units used

That is a fair picture of real sites:

Expect this mix on any list. Most small-business sites give up at least a phone number; a smaller share publish an email; a few block automated visits.

Concurrency and rate limits

Four workers is a sensible default. Your plan has a requests-per-second limit (5 on Free, 10 on Pool, more on the larger plans), enforced across all your keys. Going over it gets a 429 with a Retry-After header, which is not billed; if you raise max_workers a lot, catch that status and sleep for the number of seconds it gives. Four workers each waiting on a real website rarely get close to the limit.

For long lists, POST /v1/batch fetches up to 20 URLs in one call and bills each URL as a separate request. It suits the homepage pass; the contact-page pass depends on each homepage's links, so it stays a second step.

What it costs

1 request unit per delivered fetch: 1 or 2 per company, so 100 companies cost between 100 and 200 units. Failed fetches are not billed. See pricing.

Use the results responsibly

Collect only the contact details a business publishes for business use. Before you email or call, check the law where you and they are (GDPR and ePrivacy in the EU, for example), say who you are and why you are writing, and honour every opt-out straight away.

Next steps

Need the list of websites first? Export local business leads to CSV with Python builds one from a map search. If you would rather not run code, Lead Finder in the dashboard does the same job. The links, metadata and render options are all in the docs.

Start free with 1,000 requests Read the docs