How to track real estate listings with a daily scrape
Property portals tell you what is for sale today. They are much worse at telling you what changed since yesterday: which homes are new, which disappeared, and which just had a price cut. This guide builds that with a daily scrape of a real estate search page: one request per results page, a JSON snapshot per day, and a short diff that prints the changes.
This is a template. Real estate sites differ in markup, and many of them restrict automated access in their terms. The examples use a made-up host, listings.example, and made-up class names. Before you point it at a real portal, check its terms, find its selectors in your browser, and keep the schedule gentle.
What you are building
- A search URL with your filters already applied: area, price range, property type, sorted by newest.
- A script that reads every listing on that search (price, rooms, area, address, link) and saves them to
snapshots/2026-09-09.json. - A comparison with the previous snapshot that prints
NEW,GONEandPRICElines. - A cron entry that runs it every morning.
The scrape.land API does the fetching. It is stateless: it has no watch list and keeps no history, so the snapshots live on your disk, where you can query them however you like.
Step 1: the search page and the selectors
Run your search on the portal in a normal browser and copy the URL from the address bar. Filters are almost always in the query string, so that one URL reproduces the search. Then right-click a listing and choose Inspect. You are looking for the element that wraps one listing (here .listing-card) and, inside it, the price, the rooms, the floor area, the address and the link.
Test the selectors with one request before writing any code. A group field reads each listing from its own card, so a listing with no floor area gets null in its own row rather than shifting every area below it:
curl https://scrape.land/v1/extract \
-H "X-Api-Key: YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://listings.example/for-sale/springfield?sort=newest",
"country": "us",
"fields": {
"listings": {
"css": ".listing-card",
"fields": {
"price": ".price",
"rooms": ".rooms",
"area": ".area",
"address": ".address",
"url": {"css": "a", "attr": "href"}
}
},
"next": "a[rel=next]@href"
}}'The response has the shape {"url": …, "status": 200, "data": {"listings": [{"price": …, "rooms": …, "area": …, "address": …, "url": …}, …], "next": …}}. next is the link to the following results page, or null on the last one. If a selector is invalid, a field_errors object names it, which tells you the difference between "no listings today" and "my selector is broken".
"country" keeps every run looking at the portal from the same market, which matters on sites that change currency or listings by visitor location.
Step 2: snapshot and diff
The script keys every listing by its absolute URL. That is the one value that stays the same when the price, the photos or the description change, so it is what "the same listing" means from one day to the next.
#!/usr/bin/env python3
"""listings.py: snapshot today's listings and report new, removed and repriced ones."""
import datetime
import json
import os
import pathlib
import sys
from urllib.parse import urljoin
import requests
SEARCH = "https://listings.example/for-sale/springfield?sort=newest"
FIELDS = {
"listings": {
"css": ".listing-card",
"fields": {
"price": ".price",
"rooms": ".rooms",
"area": ".area",
"address": ".address",
"url": {"css": "a", "attr": "href"},
},
},
"next": "a[rel=next]@href",
}
SNAPS = pathlib.Path(__file__).with_name("snapshots")
def scrape():
found, url, pages = {}, SEARCH, 0
while url and pages < 10:
r = requests.post(
"https://scrape.land/v1/extract",
headers={"X-Api-Key": os.environ["SCRAPELAND_KEY"]},
json={"url": url, "country": "us", "fields": FIELDS},
timeout=90,
)
r.raise_for_status()
body = r.json()
if body.get("field_errors"):
sys.exit(f"selector broken: {body['field_errors']}")
data = body["data"]
for row in data["listings"] or []:
if row["url"]:
found[urljoin(url, row.pop("url"))] = row
pages += 1
url = urljoin(url, data["next"]) if data["next"] else None
return found
today = scrape()
if not today:
sys.exit("no listings found: check the page and selectors before trusting a diff")
SNAPS.mkdir(exist_ok=True)
stamp = datetime.date.today().isoformat()
older = sorted(p for p in SNAPS.glob("*.json") if p.stem != stamp)
(SNAPS / f"{stamp}.json").write_text(json.dumps(today, indent=1), encoding="utf-8")
if not older:
sys.exit(f"first snapshot: {len(today)} listings saved")
before = json.loads(older[-1].read_text(encoding="utf-8"))
for url in today.keys() - before.keys():
print("NEW", today[url]["price"], today[url]["address"], url)
for url in before.keys() - today.keys():
print("GONE", before[url]["price"], before[url]["address"], url)
for url in today.keys() & before.keys():
if today[url]["price"] != before[url]["price"]:
print("PRICE", before[url]["price"], "->", today[url]["price"], today[url]["address"], url)A few choices in there are deliberate:
- An empty result stops the script before anything is compared. A redesign or a blocked page would otherwise look like every listing was sold overnight.
- The price is compared as text, exactly as the site prints it. That avoids parsing currency formats, and
$350,000changing to$339,000is still a change. Parse it into a number later, when you chart it. - The page loop stops at 10 pages. A tight search is better than a huge one: split a city into neighbourhoods rather than paging through thousands of results.
- Running twice in a day overwrites today's file and still compares against yesterday's.
We checked the diff logic with two stubbed days of data. The second run printed:
NEW $299,000 3 C St https://listings.example/l/3
GONE $420,000 2 B St https://listings.example/l/2
PRICE $350,000 -> $339,000 1 A St https://listings.example/l/1Step 3: run it every morning with cron
Run crontab -e on a Linux or macOS machine and add:
SCRAPELAND_KEY=YOUR_KEY
17 7 * * * /usr/bin/python3 /opt/listings/listings.py >> /opt/listings/changes.log 2>&1That runs at 07:17 every day and appends the output to a log. Any scheduler works: a systemd timer, Windows Task Scheduler or a scheduled CI job. Once the output looks right, send the NEW and PRICE lines to email or Slack; watching a page for changes shows how to post to a Slack incoming webhook.
Making it more reliable
- Listings loaded by JavaScript. If
listingsis empty but your browser shows results, the portal builds them in the browser. Add"render": trueand"block_resources": true. See scraping JavaScript-heavy pages. - Details only on the listing page. Year built, energy rating or agent often live on the detail page. Fetch only the
NEWURLs each day rather than every listing, so the extra cost stays small. - Structured data. Many listing pages carry schema.org JSON-LD with the price and address. Fetch one with
"metadata": trueto see what is there; see page metadata and JSON-LD. - Many different portals. Writing selectors per site is the main cost. On the Scale plan and up,
"extract_type": "real_estate"returns a normalised listing object from a page without selectors.
What it costs
Each results page is 1 request unit, and you only pay for responses that land. A search that spans 3 pages, checked once a day, is about 90 units a month, well inside the Free plan's 1,000. A rendered page is 5 units, or 1 with "block_resources": true. See pricing.
Next steps
The same snapshot-and-diff pattern powers price monitoring on a schedule. The group field form and the country option are in the docs. Create a free account and your first snapshot can be on disk tomorrow morning.