Skip to content
All posts

How to export local business leads to CSV with Python

Leads and local businesses5 min read
How to export local business leads to CSV with Python: the post's first code sample

One search for "bakeries in Manchester" is easy. A list of bakeries and florists across four cities, with no duplicates, in one file, is where a short script pays off. This guide loops categories and cities through POST /v1/places, removes duplicates, and exports the local business leads to CSV with Python while keeping count of the request units spent.

What one call returns

Each call to /v1/places takes one category and one location, plus optional website, phone and email filters (each "any", "yes" or "no") and a limit. It returns a results array of businesses, count (rows in this response), matched (how many fit your filters in that area) and units (what the call cost). Rows carry name, category, address, city, lat, lon, website, profile, phone, email, opening_hours and osm_url; a field with no data is left out of the row.

So a multi-city list is a loop of calls. The only work on your side is merging the results.

The script

It needs Python 3 and requests. Set your key, your categories and your cities at the top. GET /v1/capabilities with your key lists every valid category under entitlements.places.categories.

Python
import csv
import requests

API = "https://scrape.land/v1/places"
KEY = "YOUR_KEY"
CATEGORIES = ["bakery", "florist"]
CITIES = ["Manchester, UK", "Salford, UK"]
COLS = ["name", "category", "address", "city", "phone", "email",
        "website", "profile", "opening_hours", "osm_url", "search_city"]

seen = set()
rows = []
units = 0
for city in CITIES:
    for cat in CATEGORIES:
        r = requests.post(API, headers={"X-Api-Key": KEY},
                          json={"category": cat, "location": city, "phone": "yes"},
                          timeout=60)
        if r.status_code == 502:          # map service busy, not billed
            print("busy, skipped:", cat, city)
            continue
        r.raise_for_status()
        body = r.json()
        units += body["units"]
        new = 0
        for p in body["results"]:
            key = (p["name"].strip().lower(), p.get("address", "").strip().lower())
            if key in seen:
                continue
            seen.add(key)
            p["search_city"] = city
            rows.append(p)
            new += 1
        print(f"{cat:12} {city:12} count={body['count']:3} matched={body['matched']:4} "
              f"new={new:3} units={body['units']}")

with open("leads.csv", "w", newline="", encoding="utf-8") as f:
    w = csv.DictWriter(f, fieldnames=COLS, extrasaction="ignore")
    w.writeheader()
    w.writerows(rows)

print(len(rows), "unique rows,", units, "units used")

A few choices in it are worth explaining:

A real run

With bakeries and florists in Manchester and Salford, on an account on the Free plan, the script printed:

shell
bakery       Manchester, UK count= 10 matched=  13 new= 10 units=1
florist      Manchester, UK count=  7 matched=   7 new=  7 units=1
bakery       Salford, UK  count= 10 matched=  20 new=  3 units=1
florist      Salford, UK  count= 10 matched=  13 new=  4 units=1
24 unique rows, 4 units used

Two things show up straight away. Salford sits right next to Manchester, so of its 10 bakeries only 3 were new; the other 7 had already come back in the Manchester search. That is the dedupe earning its place. And on the Free plan every search stops at 10 rows even when more matched (13 and 20 here). On any paid plan the same searches would return every match up to 200 per search.

The first lines of leads.csv:

shell
name,category,address,city,phone,email,website,profile,opening_hours,osm_url,search_city
Companio,bakery,60 Spear Street,,+44 7765 914603,,https://companiobakery.co.uk,,"Tu-Fr 08:00-14:00; Sa,Su 08:00-15:00",https://www.openstreetmap.org/node/13831823973,"Manchester, UK"
Companio Bakery,bakery,35 Radium Street,,+44 7765 914603,info@companiobakery.co.uk,https://companiobakery.co.uk,https://www.facebook.com/companiobakery,Mo off; Tu-Th 08:00-15:00; Fr-Sa 08:00-14:00; Su 08:00-12:00,https://www.openstreetmap.org/node/5571135921,"Manchester, UK"
GAIL's,bakery,"760 Wilmslow Road, M20 2DR",Manchester,+441613488783,,https://gails.com/pages/didsbury,,Mo-Fr 07:00-18:30; Sa 06:45-18:30; Su 06:45-18:00,https://www.openstreetmap.org/node/5159617444,"Manchester, UK"

The first two rows show the limit of any name-and-address key: "Companio" and "Companio Bakery" are two mapped shops with the same phone and website. They may be two branches or one business mapped twice. If you want one row per company, dedupe a second time on website or on the phone number with the spaces removed.

Watch the cost

The rule is 1 request unit per 10 rows returned, and at least 1 per search. The run above made 4 searches and returned 37 rows (24 of them unique): 4 units in total, because no single search returned more than 10 rows. The florist search in Manchester returned 7 rows and still cost 1 unit, and a search that finds nothing costs 1 unit too.

On a paid plan, where a search returns up to 200 rows, plan on up to 20 units per search. A grid of 5 categories across 10 cities is 50 searches, so at most 1,000 units. The script adds up units from each response, so you see the exact figure at the end, and GET /v1/account tells you what is left on your plan before a large run.

Two ways to spend less: filter on the server ("phone": "yes" returns fewer rows than filtering in Python after the fact, and fewer rows cost fewer units), and pass a limit when you only need the first rows. Rows come back best-contact-first: a business with a website and a phone sorts above one with neither.

Keep it clean and legal

The data is © OpenStreetMap contributors under the ODbL; keep the attribution string from the response with the file if you share it. Before you contact anyone on the list, check the rules where you and they are (in the EU, GDPR and ePrivacy apply to a sole trader's phone and email), use only public business details, and honour opt-outs.

Next steps

Many rows have a website but no email. Collecting contact details from a list of websites with Python picks up where this script stops. If you only need one category in one city, the dashboard's Local businesses tool does it without code. Plan sizes are on the pricing page, and the docs cover every endpoint.

Start free with 1,000 requests Read the docs