How to scrape news headlines into JSON
A news monitor, a research feed or a daily digest all start the same way: scrape news headlines into JSON, with the link and the time for each story, and do it again tomorrow without collecting the same stories twice. This guide does that on the Hacker News front page, then reads an article's publish date and author from its structured data.
The page and its markup
Hacker News is a good practice target: it is public, plain HTML, and changes all day, so you can see the deduplication work. Each story is a table row, tr.athing, whose id attribute is the story id. The headline link is .titleline > a. The time is in the next row, inside td.subtext, as a span.age whose title attribute holds an exact timestamp; its link points to item?id= followed by the same story id.
That split matters. A group field reads its inner fields inside each container, so one group cannot reach from the headline row into the row below it. The answer is two groups, one per row type, joined on the story id.
Scrape news headlines with two CSS groups
curl https://scrape.land/v1/extract \
-H "X-Api-Key: YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://news.ycombinator.com/",
"fields": {
"stories": {"css": "tr.athing", "fields": {
"id": {"css": ".", "attr": "id"},
"title": ".titleline > a",
"url": {"css": ".titleline > a", "attr": "href"},
"site": ".sitestr"
}},
"meta": {"css": "td.subtext", "fields": {
"item": {"css": ".age a", "attr": "href"},
"posted": {"css": ".age", "attr": "title"},
"points": ".score"
}}
}}'"css": "." inside a group means the container itself, so {"css": ".", "attr": "id"} reads the row's own id. Here is the real response from 30 September 2026, trimmed to two stories:
{
"url": "https://news.ycombinator.com/",
"status": 200,
"data": {
"stories": [
{
"id": "49913571",
"site": "blog.google",
"title": "Gemini 4 Argon",
"url": "https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/"
},
{
"id": "49912955",
"site": "quantamagazine.org",
"title": "Surprisingly complex waves reveal the brain's inner workings",
"url": "https://www.quantamagazine.org/surprisingly-complex-waves-reveal-the-brains-inner-workings-20260930/"
}
],
"meta": [
{"item": "item?id=49913571", "points": "678 points", "posted": "2026-09-30T20:04:37"},
{"item": "item?id=49912955", "points": "82 points", "posted": "2026-09-30T19:04:50"}
]
}
}Both lists had 30 entries. Notice two things. The item links are relative, exactly as the page writes them, so resolve them against the page URL. And one row on the page, a job ad, had no score: its points came back null in its own row, which is the point of using groups instead of parallel lists.
Join, dedupe and save
The script below joins the two lists on the story id, skips any story it saved on a previous run, and appends new ones to a JSON Lines file (one JSON object per line), which is easy to append to and to load later.
import json
import pathlib
from urllib.parse import urljoin
import requests
API = "https://scrape.land/v1/extract"
KEY = "YOUR_KEY"
PAGE = "https://news.ycombinator.com/"
FIELDS = {
"stories": {"css": "tr.athing", "fields": {
"id": {"css": ".", "attr": "id"},
"title": ".titleline > a",
"url": {"css": ".titleline > a", "attr": "href"},
}},
"meta": {"css": "td.subtext", "fields": {
"item": {"css": ".age a", "attr": "href"},
"posted": {"css": ".age", "attr": "title"},
}},
}
SEEN = pathlib.Path("seen.json")
r = requests.post(API, headers={"X-Api-Key": KEY},
json={"url": PAGE, "fields": FIELDS}, timeout=60)
r.raise_for_status()
data = r.json()["data"]
# Join the two lists on the story id, never on position.
posted = {m["item"].split("=")[-1]: m["posted"] for m in data["meta"] or [] if m["item"]}
seen = set(json.loads(SEEN.read_text())) if SEEN.exists() else set()
new = []
for s in data["stories"] or []:
if s["id"] in seen:
continue
new.append({
"id": s["id"],
"title": s["title"],
"url": urljoin(PAGE, s["url"] or ""),
"posted": posted.get(s["id"]),
})
SEEN.write_text(json.dumps(sorted(seen | {n["id"] for n in new})))
with open("headlines.jsonl", "a", encoding="utf-8") as f:
for n in new:
f.write(json.dumps(n, ensure_ascii=False) + "\n")
print(f"{len(new)} new of {len(data['stories'] or [])} on the page")We ran it twice in a row. The first run printed 30 new of 30 on the page; the second printed 0 new of 30 on the page. Every story got its timestamp from the join. The first lines of headlines.jsonl:
{"id": "49913571", "title": "Gemini 4 Argon", "url": "https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/", "posted": "2026-09-30T20:04:37"}
{"id": "49915082", "title": "The top secret URSALA, RAQUEL, and FARRAH satellites", "url": "https://www.thespacereview.com/article/4951/1", "posted": "2026-09-30T22:03:04"}Dedupe on a stable id when the site has one, as here. When it does not, use the article URL. Titles are a poor key: editors change them after publishing. The posted value has no timezone; on this site it is UTC, but check that on any site you scrape.
Article details from JSON-LD
A headline list rarely carries the author or an exact publish date. The article page usually does, in schema.org JSON-LD that news sites publish for search engines. POST /v1/fetch with "metadata": true returns it already parsed, along with the title, description, canonical URL and OpenGraph tags. "format": "text" keeps the rest of the response small.
curl https://scrape.land/v1/fetch \
-H "X-Api-Key: YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://www.quantamagazine.org/surprisingly-complex-waves-reveal-the-brains-inner-workings-20260930/",
"format": "text",
"metadata": true}'From the real response, the part of metadata that matters here (the text and several other fields are left out):
{
"canonical": "https://www.quantamagazine.org/surprisingly-complex-waves-reveal-the-brains-inner-workings-20260930/",
"description": "Unexpected patterns traveling across the human brain may be reorganizing its activity in real time.",
"jsonld": [
{
"@context": "https://schema.org",
"@graph": [
{"@type": "WebSite", "name": "Quanta Magazine", …},
{
"@type": "WebPage",
"author": {"@type": "Person", "name": "Conor Feehly", …},
"datePublished": "2026-09-30T14:52:12+00:00",
"dateModified": "2026-09-30T14:52:44+00:00",
"name": "Surprisingly Complex Waves Reveal the Brain’s Inner Workings | Quanta Magazine",
…
}
]
}
],
"og": {"type": "article", "site_name": "Quanta Magazine", …},
"title": "Surprisingly Complex Waves Reveal the Brain’s Inner Workings | Quanta Magazine"
}This site nests its objects inside an @graph array, and uses WebPage rather than NewsArticle. Other sites put a single NewsArticle at the top level. A small helper copes with both:
def article_details(metadata):
nodes = []
for block in metadata.get("jsonld", []):
nodes += block.get("@graph", [block])
for node in nodes:
if node.get("datePublished"):
author = node.get("author") or {}
if isinstance(author, list):
author = author[0] if author else {}
return {"published": node["datePublished"], "author": author.get("name")}
return {"published": None, "author": None}Only fetch details for stories that are new on this run. That keeps the cost proportional to the news, not to how often you check.
Adapting it to other news sites
- Most sites are simpler than this one. A typical news or blog index puts each story in one
articleor card element, with the headline, link and<time datetime>inside it. Then a single group does the whole job, and no join is needed. - Relative times. If a site only prints "2 hours ago" and has no
datetimeortitleattribute, record the time of your run next to it, or take the exact date from the article's JSON-LD. - Headlines loaded by JavaScript. If the groups come back empty, add
"render": trueand"block_resources": true. Many sites also publish an RSS feed, which is worth checking first.
What it costs
The headline page is 1 request unit per run, and each article fetch is 1 more. Checking the front page every hour is about 720 units a month; adding details for 50 new stories a day adds about 1,500. You only pay for responses that land. See pricing.
Next steps
For more on the metadata option, read page metadata and JSON-LD. To be told when a page changes rather than collecting every story, see watching a web page for changes. The group field form is in the docs, and you can sign up free to run the script above.