How to collect contact details from websites in Python
You have a column of company websites, perhaps from a map search or a trade directory, and you need an email and a phone number for each. This guide builds a small Python script that collects contact details from a list of websites: it fetches each homepage through POST /v1/fetch, falls back to the contact page when it has to, and writes a CSV. It is the same approach the dashboard's Lead Finder uses.
The plan
For every site:
- Fetch the homepage as HTML with
"links": true, which adds alinksarray of every link on the page as an absolute URL, and"metadata": true, which adds ametadataobject with the pagetitle(used as the company name). - Parse the HTML for
mailto:links, plain addresses in the text, Cloudflare-protected addresses,tel:links, the schema.orgtelephonefield and social profile links. - If an email or a phone is still missing, pick one contact or about page on the same host from
links, fetch it and parse it the same way.
That is one or two fetches per company. Each delivered fetch is 1 request unit; a fetch that fails is not billed.
What a fetch returns
A request with both options looks like this:
curl https://scrape.land/v1/fetch \
-H "X-Api-Key: YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://www.greenhalghs.com/",
"format": "html", "metadata": true, "links": true}'The response has url, status and html, plus metadata and links. For this bakery's homepage the title came back as "Greenhalgh's Craft Bakery – Fresh Breads, Pies & Cakes", and among the links were https://www.greenhalghs.com/contact-us/ and https://www.greenhalghs.com/about-us-greenhalghs-craft-bakery/. The script prefers "contact" over "about" when both exist, because the about page rarely carries an email.
The script
Python 3.9 or newer, plus requests. It reads four sites at a time with a thread pool.
import csv, re
from concurrent.futures import ThreadPoolExecutor
from urllib.parse import urlparse, unquote
import requests
API = "https://scrape.land/v1/fetch"
KEY = "YOUR_KEY"
SITES = ["https://oldroydgroup.co.uk/", "https://wbtrade.co.uk/",
"https://www.greenhalghs.com/", "https://josiescakes.co.uk/"]
EMAIL = re.compile(r"[A-Z0-9._%+-]+@[A-Z0-9.-]+\.[A-Z]{2,}", re.I)
WEBMAIL = re.compile(r"@(gmail|googlemail|yahoo|hotmail|outlook|live|aol|icloud)\.", re.I)
JUNK = re.compile(r"\.(png|jpe?g|gif|svg|webp|css|js)$|@(example|domain|sentry|wixpress)\.", re.I)
CF = re.compile(r'(?:data-cfemail=["\']|/cdn-cgi/l/email-protection#)([0-9a-f]+)', re.I)
TEL = re.compile(r'href=["\']tel:([^"\']+)["\']', re.I)
LD_TEL = re.compile(r'"telephone"\s*:\s*"([^"]+)"', re.I)
SOCIAL = re.compile(r'href=["\'](https?://(?:www\.)?(?:linkedin\.com/company|facebook\.com'
r'|instagram\.com|x\.com|twitter\.com)/[^"\'?#\s]+)', re.I)
def fetch(url, **opts):
r = requests.post(API, headers={"X-Api-Key": KEY},
json={"url": url, "format": "html", **opts}, timeout=90)
r.raise_for_status()
return r.json()
def cf_decode(hexstr):
# Cloudflare's email obfuscation: first byte is the XOR key.
key = int(hexstr[:2], 16)
return "".join(chr(int(hexstr[i:i + 2], 16) ^ key) for i in range(2, len(hexstr), 2))
def parse(html):
html = html.replace("@", "@")
emails = {unquote(m) for m in re.findall(r"mailto:([^\"'?>\s]+)", html, re.I)}
emails |= {cf_decode(h) for h in CF.findall(html)}
visible = re.sub(r"<script(?![^>]*ld\+json).*?</script>|<style.*?</style>", " ",
html, flags=re.S | re.I)
emails |= set(EMAIL.findall(visible))
emails = {e.lower() for e in emails
if EMAIL.fullmatch(e) and not WEBMAIL.search(e) and not JUNK.search(e)}
phones = {re.sub(r"\D", "", p): unquote(p).strip() # one entry per number
for p in TEL.findall(html) + LD_TEL.findall(html)}
return emails, set(phones.values()), set(SOCIAL.findall(html))
def contact_page(links, host):
same = [l for l in links if (urlparse(l).hostname or "").removeprefix("www.") == host]
for word in ("contact", "kontakt", "about", "impressum"):
for l in same:
if word in urlparse(l).path.lower():
return l
return None
def lead(site):
host = urlparse(site).hostname.removeprefix("www.")
row = {"website": site, "company": host, "emails": set(), "phones": set(),
"social": set(), "units": 0, "error": ""}
try:
d = fetch(site, metadata=True, links=True)
row["units"] += 1
title = (d.get("metadata") or {}).get("title") or ""
row["company"] = " ".join(title.split()) or host
row["emails"], row["phones"], row["social"] = parse(d.get("html", ""))
page = contact_page(d.get("links", []), host)
if page and (not row["emails"] or not row["phones"]):
e, p, s = parse(fetch(page).get("html", ""))
row["units"] += 1
row["emails"] |= e
row["phones"] |= p
row["social"] |= s
except requests.HTTPError as e:
row["error"] = f'{e.response.status_code} {e.response.json().get("error", "")}'
return row
with ThreadPoolExecutor(max_workers=4) as pool:
rows = list(pool.map(lead, SITES))
with open("contacts.csv", "w", newline="", encoding="utf-8") as f:
w = csv.writer(f)
w.writerow(["company", "website", "emails", "phones", "social", "error"])
for r in rows:
w.writerow([r["company"], r["website"], "; ".join(sorted(r["emails"])),
"; ".join(sorted(r["phones"])), "; ".join(sorted(r["social"])), r["error"]])
print(f'{r["website"]:32} emails={sorted(r["emails"])} phones={sorted(r["phones"])} '
f'social={len(r["social"])} {r["error"]}')
print(sum(r["units"] for r in rows), "request units used")Some notes on the parsing:
- Webmail is skipped. A Gmail or Hotmail address on a company site is usually one person's private inbox, so the script drops it and keeps addresses on business domains.
- Junk is skipped. Image file names such as
logo@2x.pnglook like email addresses to a regular expression; so do tracking addresses from site builders. TheJUNKpattern removes them. - Scripts are ignored for plain-text matches, except JSON-LD, where schema.org
emailandtelephonefields live. - Phones are deduplicated by their digits, so
01204 696204and01204696204on the same page become one entry.
A real run
We ran the script against four small UK business websites taken from a map search. It printed:
https://oldroydgroup.co.uk/ emails=[] phones=[] social=0
https://wbtrade.co.uk/ emails=['sales@wbtrade.co.uk'] phones=['01924267488'] social=0
https://www.greenhalghs.com/ emails=[] phones=['01204 696204'] social=3
https://josiescakes.co.uk/ emails=[] phones=[] social=0
5 request units usedThat is a fair picture of real sites:
- WB Trade hides its address behind Cloudflare's email protection, so a plain regular expression finds nothing. The
cf_decodestep recoverssales@wbtrade.co.uk. - Greenhalgh's publishes a phone and three social profiles but no email address on either its homepage or its contact page. The script read the homepage and the contact page (2 units) and correctly reports no email.
- Oldroyd and Josie's came back with a page titled "One moment, please...": a bot check, not the site. The fetch succeeded, so it counts, but there is nothing to parse. On plans with browser rendering, retry those rows with
"render": trueadded to the fetch; rendering costs more units than a plain fetch, so do it only for the rows that need it. On an earlier run one of these sites answered502with "could not fetch the target page", which the script records in theerrorcolumn and which was not billed.
Expect this mix on any list. Most small-business sites give up at least a phone number; a smaller share publish an email; a few block automated visits.
Concurrency and rate limits
Four workers is a sensible default. Your plan has a requests-per-second limit (5 on Free, 10 on Pool, more on the larger plans), enforced across all your keys. Going over it gets a 429 with a Retry-After header, which is not billed; if you raise max_workers a lot, catch that status and sleep for the number of seconds it gives. Four workers each waiting on a real website rarely get close to the limit.
For long lists, POST /v1/batch fetches up to 20 URLs in one call and bills each URL as a separate request. It suits the homepage pass; the contact-page pass depends on each homepage's links, so it stays a second step.
What it costs
1 request unit per delivered fetch: 1 or 2 per company, so 100 companies cost between 100 and 200 units. Failed fetches are not billed. See pricing.
Use the results responsibly
Collect only the contact details a business publishes for business use. Before you email or call, check the law where you and they are (GDPR and ePrivacy in the EU, for example), say who you are and why you are writing, and honour every opt-out straight away.
Next steps
Need the list of websites first? Export local business leads to CSV with Python builds one from a map search. If you would rather not run code, Lead Finder in the dashboard does the same job. The links, metadata and render options are all in the docs.