Skip to content
All posts

How to scrape news headlines into JSON

Use cases6 min read
How to scrape news headlines into JSON: the post's first code sample

A news monitor, a research feed or a daily digest all start the same way: scrape news headlines into JSON, with the link and the time for each story, and do it again tomorrow without collecting the same stories twice. This guide does that on the Hacker News front page, then reads an article's publish date and author from its structured data.

The page and its markup

Hacker News is a good practice target: it is public, plain HTML, and changes all day, so you can see the deduplication work. Each story is a table row, tr.athing, whose id attribute is the story id. The headline link is .titleline > a. The time is in the next row, inside td.subtext, as a span.age whose title attribute holds an exact timestamp; its link points to item?id= followed by the same story id.

That split matters. A group field reads its inner fields inside each container, so one group cannot reach from the headline row into the row below it. The answer is two groups, one per row type, joined on the story id.

Scrape news headlines with two CSS groups

curl
curl https://scrape.land/v1/extract \
  -H "X-Api-Key: YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://news.ycombinator.com/",
       "fields": {
         "stories": {"css": "tr.athing", "fields": {
           "id":    {"css": ".", "attr": "id"},
           "title": ".titleline > a",
           "url":   {"css": ".titleline > a", "attr": "href"},
           "site":  ".sitestr"
         }},
         "meta": {"css": "td.subtext", "fields": {
           "item":   {"css": ".age a", "attr": "href"},
           "posted": {"css": ".age", "attr": "title"},
           "points": ".score"
         }}
       }}'

"css": "." inside a group means the container itself, so {"css": ".", "attr": "id"} reads the row's own id. Here is the real response from 30 September 2026, trimmed to two stories:

JSON
{
  "url": "https://news.ycombinator.com/",
  "status": 200,
  "data": {
    "stories": [
      {
        "id": "49913571",
        "site": "blog.google",
        "title": "Gemini 4 Argon",
        "url": "https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/"
      },
      {
        "id": "49912955",
        "site": "quantamagazine.org",
        "title": "Surprisingly complex waves reveal the brain's inner workings",
        "url": "https://www.quantamagazine.org/surprisingly-complex-waves-reveal-the-brains-inner-workings-20260930/"
      }
    ],
    "meta": [
      {"item": "item?id=49913571", "points": "678 points", "posted": "2026-09-30T20:04:37"},
      {"item": "item?id=49912955", "points": "82 points", "posted": "2026-09-30T19:04:50"}
    ]
  }
}

Both lists had 30 entries. Notice two things. The item links are relative, exactly as the page writes them, so resolve them against the page URL. And one row on the page, a job ad, had no score: its points came back null in its own row, which is the point of using groups instead of parallel lists.

Join, dedupe and save

The script below joins the two lists on the story id, skips any story it saved on a previous run, and appends new ones to a JSON Lines file (one JSON object per line), which is easy to append to and to load later.

Python
import json
import pathlib
from urllib.parse import urljoin
import requests

API = "https://scrape.land/v1/extract"
KEY = "YOUR_KEY"
PAGE = "https://news.ycombinator.com/"
FIELDS = {
    "stories": {"css": "tr.athing", "fields": {
        "id": {"css": ".", "attr": "id"},
        "title": ".titleline > a",
        "url": {"css": ".titleline > a", "attr": "href"},
    }},
    "meta": {"css": "td.subtext", "fields": {
        "item": {"css": ".age a", "attr": "href"},
        "posted": {"css": ".age", "attr": "title"},
    }},
}
SEEN = pathlib.Path("seen.json")

r = requests.post(API, headers={"X-Api-Key": KEY},
                  json={"url": PAGE, "fields": FIELDS}, timeout=60)
r.raise_for_status()
data = r.json()["data"]

# Join the two lists on the story id, never on position.
posted = {m["item"].split("=")[-1]: m["posted"] for m in data["meta"] or [] if m["item"]}
seen = set(json.loads(SEEN.read_text())) if SEEN.exists() else set()

new = []
for s in data["stories"] or []:
    if s["id"] in seen:
        continue
    new.append({
        "id": s["id"],
        "title": s["title"],
        "url": urljoin(PAGE, s["url"] or ""),
        "posted": posted.get(s["id"]),
    })

SEEN.write_text(json.dumps(sorted(seen | {n["id"] for n in new})))
with open("headlines.jsonl", "a", encoding="utf-8") as f:
    for n in new:
        f.write(json.dumps(n, ensure_ascii=False) + "\n")
print(f"{len(new)} new of {len(data['stories'] or [])} on the page")

We ran it twice in a row. The first run printed 30 new of 30 on the page; the second printed 0 new of 30 on the page. Every story got its timestamp from the join. The first lines of headlines.jsonl:

JSON
{"id": "49913571", "title": "Gemini 4 Argon", "url": "https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/", "posted": "2026-09-30T20:04:37"}
{"id": "49915082", "title": "The top secret URSALA, RAQUEL, and FARRAH satellites", "url": "https://www.thespacereview.com/article/4951/1", "posted": "2026-09-30T22:03:04"}

Dedupe on a stable id when the site has one, as here. When it does not, use the article URL. Titles are a poor key: editors change them after publishing. The posted value has no timezone; on this site it is UTC, but check that on any site you scrape.

Article details from JSON-LD

A headline list rarely carries the author or an exact publish date. The article page usually does, in schema.org JSON-LD that news sites publish for search engines. POST /v1/fetch with "metadata": true returns it already parsed, along with the title, description, canonical URL and OpenGraph tags. "format": "text" keeps the rest of the response small.

curl
curl https://scrape.land/v1/fetch \
  -H "X-Api-Key: YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://www.quantamagazine.org/surprisingly-complex-waves-reveal-the-brains-inner-workings-20260930/",
       "format": "text",
       "metadata": true}'

From the real response, the part of metadata that matters here (the text and several other fields are left out):

JSON
{
  "canonical": "https://www.quantamagazine.org/surprisingly-complex-waves-reveal-the-brains-inner-workings-20260930/",
  "description": "Unexpected patterns traveling across the human brain may be reorganizing its activity in real time.",
  "jsonld": [
    {
      "@context": "https://schema.org",
      "@graph": [
        {"@type": "WebSite", "name": "Quanta Magazine", …},
        {
          "@type": "WebPage",
          "author": {"@type": "Person", "name": "Conor Feehly", …},
          "datePublished": "2026-09-30T14:52:12+00:00",
          "dateModified": "2026-09-30T14:52:44+00:00",
          "name": "Surprisingly Complex Waves Reveal the Brain’s Inner Workings | Quanta Magazine",
          …
        }
      ]
    }
  ],
  "og": {"type": "article", "site_name": "Quanta Magazine", …},
  "title": "Surprisingly Complex Waves Reveal the Brain’s Inner Workings | Quanta Magazine"
}

This site nests its objects inside an @graph array, and uses WebPage rather than NewsArticle. Other sites put a single NewsArticle at the top level. A small helper copes with both:

Python
def article_details(metadata):
    nodes = []
    for block in metadata.get("jsonld", []):
        nodes += block.get("@graph", [block])
    for node in nodes:
        if node.get("datePublished"):
            author = node.get("author") or {}
            if isinstance(author, list):
                author = author[0] if author else {}
            return {"published": node["datePublished"], "author": author.get("name")}
    return {"published": None, "author": None}

Only fetch details for stories that are new on this run. That keeps the cost proportional to the news, not to how often you check.

Adapting it to other news sites

What it costs

The headline page is 1 request unit per run, and each article fetch is 1 more. Checking the front page every hour is about 720 units a month; adding details for 50 new stories a day adds about 1,500. You only pay for responses that land. See pricing.

Next steps

For more on the metadata option, read page metadata and JSON-LD. To be told when a page changes rather than collecting every story, see watching a web page for changes. The group field form is in the docs, and you can sign up free to run the script above.

Start free with 1,000 requests Read the docs