Skip to content
All posts

How to scrape event listings into a calendar

Use cases5 min read
How to scrape event listings into a calendar: the post's first code sample

Conference lists, venue programmes and meetup pages all show events as text, and none of them adds anything to your calendar. This guide shows how to scrape event listings into a calendar: pull each event's title, date, venue and link from a listings page, turn the dates into real dates, and write an .ics file that Google Calendar, Outlook and Apple Calendar can import. The example is the Python events list on python.org.

Look at the page first

Open python.org/events/python-events and inspect one event. Each is an li in ul.list-recent-events, with the title link in an h3, a <time> element and a .event-location. The time element is the useful part: its datetime attribute holds the start date in ISO format, so you do not have to parse "07 Oct." for it.

Extract every event

A group reads the same fields inside each event, so every row stays together even when one event has no venue:

curl
curl https://scrape.land/v1/extract \
  -H "X-Api-Key: YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://www.python.org/events/python-events/",
       "fields": {
         "events": {
           "css": "ul.list-recent-events li",
           "fields": {
             "title": "h3 a",
             "url":   {"css": "h3 a", "attr": "href"},
             "start": {"css": "time", "attr": "datetime"},
             "when":  "time",
             "venue": ".event-location"
           }
         }
       }}'

The real response, first two of six events shown, with the whitespace inside when shortened:

JSON
{
  "url": "https://www.python.org/events/python-events/",
  "status": 200,
  "data": {
    "events": [
      {
        "title": "PyBay 2026",
        "url": "/events/python-events/2136/",
        "start": "2026-10-03T00:00:00+00:00",
        "when": "03 Oct.\n        \n            2026 …",
        "venue": "San Francisco, CA, USA"
      },
      {
        "title": "PyCon Africa 2026",
        "url": "/events/python-events/2187/",
        "start": "2026-10-07T00:00:00+00:00",
        "when": "07 Oct.\n        \n            2026 … –\n            11 Oct. …",
        "venue": "Kampala, Uganda"
      }
    ]
  }
}

Two things to handle. The url is relative, so it needs the page's address in front of it. And the end date only appears in the text of when: a one-day event has a single date, a multi-day event has "07 Oct. 2026 – 11 Oct. 2026" once the whitespace is collapsed.

Parse the dates and write the .ics file

An .ics file is plain text: a VCALENDAR holding one VEVENT per event. These events are all-day, so they use VALUE=DATE, and the end date is exclusive: an event on 3 October ends on 4 October. The script needs only requests:

Python
import re
from datetime import date, datetime, timedelta, timezone
from urllib.parse import urljoin
import requests

PAGE = "https://www.python.org/events/python-events/"
FIELDS = {
    "events": {
        "css": "ul.list-recent-events li",
        "fields": {
            "title": "h3 a",
            "url": {"css": "h3 a", "attr": "href"},
            "start": {"css": "time", "attr": "datetime"},
            "when": "time",
            "venue": ".event-location",
        },
    }
}

r = requests.post("https://scrape.land/v1/extract", headers={"X-Api-Key": "YOUR_KEY"},
                  json={"url": PAGE, "fields": FIELDS}, timeout=60)
r.raise_for_status()
events = r.json()["data"]["events"] or []

def end_date(start, when):
    """'07 Oct. 2026 – 11 Oct. 2026' -> 2026-10-11; one-day events end on the start."""
    m = re.search(r"–\s*(\d{1,2} [A-Z][a-z]{2})", " ".join((when or "").split()))
    if not m:
        return start
    end = datetime.strptime(f"{m.group(1)} {start.year}", "%d %b %Y").date()
    return end if end >= start else end.replace(year=start.year + 1)

def esc(text):
    """Escape a value for an iCalendar TEXT property."""
    return (text or "").replace("\\", "\\\\").replace(";", "\\;").replace(",", "\\,").replace("\n", "\\n")

stamp = datetime.now(timezone.utc).strftime("%Y%m%dT%H%M%SZ")
lines = ["BEGIN:VCALENDAR", "VERSION:2.0", "PRODID:-//events-to-ics//EN"]
for e in events:
    if not e["start"]:
        continue
    start = date.fromisoformat(e["start"][:10])
    end = end_date(start, e["when"]) + timedelta(days=1)  # DTEND is exclusive
    link = urljoin(PAGE, e["url"] or "")
    lines += [
        "BEGIN:VEVENT",
        f"UID:{link}",
        f"DTSTAMP:{stamp}",
        f"DTSTART;VALUE=DATE:{start:%Y%m%d}",
        f"DTEND;VALUE=DATE:{end:%Y%m%d}",
        f"SUMMARY:{esc(e['title'])}",
        f"LOCATION:{esc(' '.join((e['venue'] or '').split()))}",
        f"URL:{link}",
        "END:VEVENT",
    ]
lines.append("END:VCALENDAR")

with open("events.ics", "w", newline="", encoding="utf-8") as f:
    f.write("\r\n".join(lines) + "\r\n")
print(len(events), "events written to events.ics")

Details that matter to calendar apps:

Running it printed 6 events written to events.ics. The start of the real file:

shell
BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//events-to-ics//EN
BEGIN:VEVENT
UID:https://www.python.org/events/python-events/2136/
DTSTAMP:20260930T221141Z
DTSTART;VALUE=DATE:20261003
DTEND;VALUE=DATE:20261004
SUMMARY:PyBay 2026
LOCATION:San Francisco\, CA\, USA
URL:https://www.python.org/events/python-events/2136/
END:VEVENT
BEGIN:VEVENT
UID:https://www.python.org/events/python-events/2187/
DTSTAMP:20260930T221141Z
DTSTART;VALUE=DATE:20261007
DTEND;VALUE=DATE:20261012
SUMMARY:PyCon Africa 2026
LOCATION:Kampala\, Uganda
URL:https://www.python.org/events/python-events/2187/
END:VEVENT

You rarely want every event. Filter the list before writing it: keep only rows whose venue contains a country you can travel to, or whose title contains a keyword, or drop events that start more than 90 days out. The filter is one if at the top of the loop, and a smaller calendar is one you will actually read.

PyCon Africa runs 7 to 11 October, so its DTEND is the 12th. Import events.ics in Google Calendar (Settings, Import and export), Outlook or Apple Calendar. To have a calendar that stays current, run the script on a schedule and publish the file at a URL your calendar app subscribes to.

Pages that publish Event JSON-LD

Many ticketing and venue sites describe each event in schema.org JSON-LD for search engines, with startDate, endDate, location and url in fixed formats. When a page has it, you can skip the selectors and the date parsing. Fetch with "metadata": true and filter metadata.jsonld for events. The host below is illustrative:

Python
import requests

r = requests.post("https://scrape.land/v1/fetch",
                  headers={"X-Api-Key": "YOUR_KEY"},
                  json={"url": "https://venue.example/whats-on", "format": "text", "metadata": True},
                  timeout=60)
r.raise_for_status()
for node in r.json().get("metadata", {}).get("jsonld", []):
    if isinstance(node, dict) and str(node.get("@type", "")).endswith("Event"):
        place = node.get("location") or {}
        print(node.get("name"), node.get("startDate"), node.get("endDate"),
              place.get("name") if isinstance(place, dict) else place, node.get("url"))

The endswith("Event") check also catches subtypes such as MusicEvent and SportsEvent. JSON-LD dates often carry a time and timezone (2026-11-14T19:30:00+01:00); for timed events, write DTSTART in UTC with a trailing Z instead of VALUE=DATE. Not every site has it: the python.org event page we checked carries only a WebSite block, which is why this guide reads the HTML.

What it costs

The listings page above is one CSS extraction: 1 request unit for all six events. A fetch with metadata is also 1 unit. Blocked or failed requests are not billed. See pricing.

Next steps

More on reading structured data is in page metadata and JSON-LD. If the listings spread over several pages, see scraping a paginated catalog; the same loop works for events. Every field option is in the docs, and you can create a free account to run the script.

Start free with 1,000 requests Read the docs