Skip to content
All posts

How to watch a web page for changes and get alerted

Use cases5 min read
How to watch a web page for changes and get alerted: the post's first code sample

A back-in-stock notice, a changed policy page, a new version on a download page: often you only need to know that a page changed, and what changed. This guide shows how to watch a web page for changes with a short script: fetch the part you care about, hash it, compare it with the last run, and send an alert to Slack, a webhook or email when it differs.

Pick what to watch

There are two ways to decide what counts as a change, and the choice matters more than any code:

Both are one request, and you hash the result. A hash is a short fingerprint: the same text always gives the same hash, and any change gives a different one, so you only need to store one string per page.

Watching one field

Here we watch the price and stock line of a book on the practice shop Books to Scrape:

curl
curl https://scrape.land/v1/extract \
  -H "X-Api-Key: YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html",
       "fields": {"price": ".product_main .price_color",
                  "stock": ".product_main .availability"}}'
JSON
{
  "url": "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html",
  "status": 200,
  "data": {
    "price": "£51.77",
    "stock": "In stock (22 available)"
  }
}

The .product_main prefix scopes the selectors to the main product. This page has only one price, but most real shop pages also list related products with the same price classes, and a field returns the first match on the page, which may not be the one you meant.

The watch script

Python
#!/usr/bin/env python3
"""watch.py: alert when the watched part of a page changes."""
import hashlib
import json
import os
import pathlib
import sys
import requests

URL = "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html"
FIELDS = {"price": ".product_main .price_color", "stock": ".product_main .availability"}
STATE = pathlib.Path(__file__).with_name("watch-state.json")
SLACK = os.environ.get("SLACK_WEBHOOK_URL")

def notify(text):
    print(text)
    if SLACK:
        requests.post(SLACK, json={"text": text}, timeout=15).raise_for_status()

r = requests.post(
    "https://scrape.land/v1/extract",
    headers={"X-Api-Key": os.environ["SCRAPELAND_KEY"]},
    json={"url": URL, "fields": FIELDS},
    timeout=90,
)
r.raise_for_status()
body = r.json()
now = body["data"]
if body["status"] != 200 or body.get("field_errors") or not now.get("price"):
    sys.exit(f"check failed, not comparing: status {body['status']}, {body.get('field_errors') or now}")

digest = hashlib.sha256(json.dumps(now, sort_keys=True).encode()).hexdigest()
old = json.loads(STATE.read_text()) if STATE.exists() else None
STATE.write_text(json.dumps({"hash": digest, "data": now}))

if old is None:
    print("first run, saved", now)
elif old["hash"] != digest:
    notify(f"Changed: {URL}\nwas: {old['data']}\nnow: {now}")
else:
    print("no change", digest[:12])

The most important line is the check before the comparison. A watcher that treats a failed fetch as a change will wake you up for nothing: a 404, a broken selector (reported in field_errors) or an empty value means "could not check", not "the page changed". The script exits with an error instead, and cron's log shows it. It also stores the data next to the hash, so the alert can say what it was and what it is now.

We ran it against the real page. The first run printed first run, saved {'price': '£51.77', 'stock': 'In stock (22 available)'} and the second printed no change 3bc10b56ace5. To see an alert without waiting for the shop, we edited the saved stock line in watch-state.json by hand and ran it again:

shell
Changed: https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html
was: {'price': '£51.77', 'stock': 'In stock (23 available)'}
now: {'price': '£51.77', 'stock': 'In stock (22 available)'}

Watching a whole page as Markdown

To watch everything on a page, swap the extraction for a Markdown fetch and hash the markdown string:

curl
curl https://scrape.land/v1/fetch \
  -H "X-Api-Key: YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://quotes.toscrape.com/", "format": "markdown"}'
JSON
{
  "url": "https://quotes.toscrape.com/",
  "status": 200,
  "markdown": "# [Quotes to Scrape](https://quotes.toscrape.com/)\n\n[Login](https://quotes.toscrape.com/login)\n\n“The world as we have created it is a process of our thinking. …"
}

We fetched it twice: both times the Markdown was 4,239 characters with the same SHA-256 hash, 7bb4e663…, so a static page does not trigger false alerts. Real pages are noisier. A "last updated" timestamp, a visitor counter or a rotating promotion will change the hash on every run. If that happens, go back to watching a field, or strip the noisy lines from the Markdown before hashing.

If the part you care about is built by JavaScript, add "render": true and "block_resources": true to either request.

Where to send the alert

Python
import os
import smtplib
from email.message import EmailMessage

def notify_email(text):
    msg = EmailMessage()
    msg["Subject"], msg["From"], msg["To"] = "Page changed", "watch@yourdomain.example", "you@yourdomain.example"
    msg.set_content(text)
    with smtplib.SMTP_SSL("smtp.yourdomain.example", 465) as s:
        s.login("watch@yourdomain.example", os.environ["SMTP_PASSWORD"])
        s.send_message(msg)

Run it on a schedule

Run crontab -e on a Linux or macOS machine and add:

crontab
SCRAPELAND_KEY=YOUR_KEY
SLACK_WEBHOOK_URL=https://hooks.slack.com/services/YOUR/WEBHOOK/URL
*/30 * * * * /usr/bin/python3 /opt/watch/watch.py >> /opt/watch/watch.log 2>&1

That checks every 30 minutes. Match the interval to how fast you need to know: hourly or daily is plenty for most pages, and it is kinder to the site you are watching. To watch several pages, give each one its own state file, or keep a dictionary of URL to hash in one file.

What it costs

Each check is 1 request unit, whether it reads fields or fetches Markdown, and you only pay for responses that land. One page checked every 30 minutes is about 1,440 units a month; checked hourly, about 720, inside the Free plan's 1,000. A rendered check is 5 units, or 1 with "block_resources": true. See pricing.

Next steps

For a price history rather than a single alert, see monitoring competitor prices on a schedule. The Markdown output is described in clean Markdown for LLMs and RAG and in the docs. Sign up free and your first watch can be running in a few minutes.

Start free with 1,000 requests Read the docs