How to scrape job listings into a spreadsheet
Recruiters, job seekers and labour-market researchers all end up wanting the same thing: a spreadsheet with one row per job, holding the title, the company, the location and a link. This guide scrapes job listings into a CSV with one CSS group request per page, follows the pagination until it runs out, and runs for real against a public practice job board.
The practice page
Real Python publishes a static job board for scraping practice at realpython.github.io/fake-jobs. It lists 100 invented jobs, each in a card with a title, a company, a location, a posting date and two links, "Learn" and "Apply". The layout is the same as most real boards: a repeating card, and the same few elements inside each one. Everything below works on a real board once you swap in its selectors.
Open the page, right-click a job title and choose Inspect. Each job sits in a div.card-content. Inside it are h2.title, h3.company, p.location and a <time datetime="…"> element for the date.
One request per page, one row per job
Send those selectors to POST /v1/extract as a group: an object with a css selector that picks the repeating container, plus its own fields, which are read inside each container. You get back a list of objects, one per job. If one card is missing its location, that row gets null for the location instead of every location below it moving up a row, which is what happens when you collect titles and locations as two separate lists.
curl https://scrape.land/v1/extract \
-H "X-Api-Key: YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://realpython.github.io/fake-jobs/",
"fields": {
"jobs": {
"css": ".card-content",
"fields": {
"title": "h2.title",
"company": "h3.company",
"location": "p.location",
"posted": "time@datetime",
"apply": {"css": "a:contains(\"Apply\")", "attr": "href"}
}
}
}}'This is the real response, trimmed to the first three of the 100 jobs:
{
"url": "https://realpython.github.io/fake-jobs/",
"status": 200,
"data": {
"jobs": [
{
"apply": "https://realpython.github.io/fake-jobs/jobs/senior-python-developer-0.html",
"company": "Payne, Roberts and Davis",
"location": "Stewartbury, AA",
"posted": "2021-04-08",
"title": "Senior Python Developer"
},
{
"apply": "https://realpython.github.io/fake-jobs/jobs/energy-engineer-1.html",
"company": "Vasquez-Davidson",
"location": "Christopherville, AA",
"posted": "2021-04-08",
"title": "Energy engineer"
},
{
"apply": "https://realpython.github.io/fake-jobs/jobs/legal-executive-2.html",
"company": "Jackson, Chambers and Levy",
"location": "Port Ericaburgh, AA",
"posted": "2021-04-08",
"title": "Legal executive"
}
]
}
}Three details in that request are worth copying:
time@datetimereads thedatetimeattribute instead of the visible text. Job boards often print "3 days ago" but keep an exact date in an attribute. Read the attribute whenever there is one.a:contains("Apply")picks the link whose text contains "Apply".:contains()is an extension of the selector engine we use, not standard CSS, so it will not work in your browser's devtools. It is handy when two links share the same classes.- Whitespace around the text is trimmed for you: the location on the page is padded with spaces and newlines, and it arrives as
"Stewartbury, AA".
Follow the pages and write the CSV
Most boards spread results over several pages. The API fetches exactly the URL you send and does not follow links for you, so the loop lives in your script: read the jobs on a page, read the link to the next page, repeat until there is no next link. Ask for the next link in the same request by adding one more field, so each page costs one call.
import csv
from urllib.parse import urljoin
import requests
API = "https://scrape.land/v1/extract"
KEY = "YOUR_KEY"
START = "https://realpython.github.io/fake-jobs/"
FIELDS = {
"jobs": {
"css": ".card-content",
"fields": {
"title": "h2.title",
"company": "h3.company",
"location": "p.location",
"posted": "time@datetime",
"apply": {"css": 'a:contains("Apply")', "attr": "href"},
},
},
"next": "a[rel=next]@href",
}
rows, url, pages = [], START, 0
while url and pages < 20: # hard stop, in case "next" ever loops
r = requests.post(API, headers={"X-Api-Key": KEY},
json={"url": url, "fields": FIELDS}, timeout=60)
r.raise_for_status()
body = r.json()
if body.get("field_errors"):
raise SystemExit(f"selector problem: {body['field_errors']}")
data = body["data"]
for job in data["jobs"] or []:
job["apply"] = urljoin(url, job["apply"] or "")
rows.append(job)
pages += 1
url = urljoin(url, data["next"]) if data["next"] else None
with open("jobs.csv", "w", newline="", encoding="utf-8") as f:
w = csv.DictWriter(f, fieldnames=["title", "company", "location", "posted", "apply"])
w.writeheader()
w.writerows(rows)
print(f"{len(rows)} jobs from {pages} page(s) written to jobs.csv")The practice board fits on one page, so next comes back null on the first call and the loop stops. Our run printed 100 jobs from 1 page(s) written to jobs.csv, and the file starts like this:
title,company,location,posted,apply
Senior Python Developer,"Payne, Roberts and Davis","Stewartbury, AA",2021-04-08,https://realpython.github.io/fake-jobs/jobs/senior-python-developer-0.html
Energy engineer,Vasquez-Davidson,"Christopherville, AA",2021-04-08,https://realpython.github.io/fake-jobs/jobs/energy-engineer-1.htmlOn a real board, point next at its own next-page link. a[rel=next] is a common convention; other boards use something like .pagination .next a. Links are returned exactly as the page writes them, often relative (/jobs?page=2), which is why the script passes them through urljoin. If a board numbers its pages in the URL instead (?page=2, ?page=3), you can loop over the numbers and stop at the first page that returns no jobs. For many known page URLs at once, batch them in one call.
Adding the job description
The listing page rarely carries the full description. Once you have the links, fetch each job page with a flat field map. On the practice site the detail page has a site header h1 before the job's own h1, so the selectors are scoped to the job box:
curl https://scrape.land/v1/extract \
-H "X-Api-Key: YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://realpython.github.io/fake-jobs/jobs/senior-python-developer-0.html",
"fields": {"title": ".box h1", "description": ".box .content p"}}'{
"url": "https://realpython.github.io/fake-jobs/jobs/senior-python-developer-0.html",
"status": 200,
"data": {
"title": "Senior Python Developer",
"description": "Professional asset web application environmentally friendly detail-oriented asset. Coordinate educational dashboard agile employ growth opportunity. …"
}
}Our first try used h1.title and got "Fake Python", the site name, because a field returns the first match. When a value looks wrong, the fix is almost always a more specific selector.
Things that trip up job scrapers
- Duplicates across pages. New jobs are posted while you page through, which pushes older ones onto the next page. Dedupe on the job link before you write the file.
- Jobs loaded by JavaScript. If
jobscomes back as an empty list but you can see jobs in your browser, the board builds them in the browser. Add"render": true, and"block_resources": trueto keep it cheap. See scraping JavaScript-heavy pages. - Location-dependent results. Some boards show different jobs by visitor country. Add
"country": "us"(or your market) so every page is seen from the same place. - Terms of use. Read the board's terms before you scrape it, keep your request rate modest, and do not republish listings you have no right to.
What it costs
Each page is 1 request unit on every plan, and you only pay for responses that land. The practice board above is 1 unit for 100 jobs; a board with 30 pages of results is 30 units per full sweep. See pricing.
Next steps
The group form and every other field form are in the docs. For the same pattern on a shop, read scraping a paginated product catalog, and to land the rows in a shared sheet, see pulling scraped data into Google Sheets. Sign up free to get a key.