How to scrape product reviews and ratings
Reviews tell you what customers like and hate about a product, and star ratings let you compare a whole category at a glance. This guide shows how to scrape product reviews and ratings: star ratings that are hidden in class names, review blocks with an author, a rating, a date and text, the summary rating many shops publish as JSON-LD, and the pagination that reviews almost always come with.
Star ratings live in attributes, not text
A star rating is usually drawn with icons, so there is no "4" in the page text for a selector to read. The number is in the markup instead: a class name like stars-4 or star-rating Four, an attribute like data-rating="4" or aria-label="4 out of 5 stars", or a width such as style="width: 80%". Inspect one rating in your browser to see which one the site uses, then read that attribute.
The practice shop Books to Scrape uses a class: every book has a p.star-rating with a second class, One to Five. Read the class attribute inside a group, one row per book:
curl https://scrape.land/v1/extract \
-H "X-Api-Key: YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://books.toscrape.com/catalogue/category/books/mystery_3/index.html",
"fields": {
"books": {
"css": "article.product_pod",
"fields": {
"title": {"css": "h3 a", "attr": "title"},
"rating": {"css": ".star-rating", "attr": "class"},
"url": {"css": "h3 a", "attr": "href"}
}
},
"next": "li.next a@href"
}}'The real response, trimmed to the first two of 20 books on the page:
{
"url": "https://books.toscrape.com/catalogue/category/books/mystery_3/index.html",
"status": 200,
"data": {
"books": [
{
"rating": "star-rating Four",
"title": "Sharp Objects",
"url": "../../../sharp-objects_997/index.html"
},
{
"rating": "star-rating One",
"title": "In a Dark, Dark Wood",
"url": "../../../in-a-dark-dark-wood_963/index.html"
}
],
"next": "page-2.html"
}
}The title comes from the link's title attribute because the visible link text is cut short for long titles. The links are relative, exactly as the page writes them, and next points to the following page.
All pages, into numbers and a CSV
import csv
from urllib.parse import urljoin
import requests
STARS = {"One": 1, "Two": 2, "Three": 3, "Four": 4, "Five": 5}
FIELDS = {
"books": {
"css": "article.product_pod",
"fields": {
"title": {"css": "h3 a", "attr": "title"},
"rating": {"css": ".star-rating", "attr": "class"},
"url": {"css": "h3 a", "attr": "href"},
},
},
"next": "li.next a@href",
}
def stars(cls):
"""'star-rating Three' -> 3"""
return next((STARS[w] for w in (cls or "").split() if w in STARS), None)
rows, url = [], "https://books.toscrape.com/catalogue/category/books/mystery_3/index.html"
while url:
r = requests.post("https://scrape.land/v1/extract", headers={"X-Api-Key": "YOUR_KEY"},
json={"url": url, "fields": FIELDS}, timeout=60)
r.raise_for_status()
data = r.json()["data"]
for b in data["books"] or []:
rows.append({"title": b["title"], "stars": stars(b["rating"]), "url": urljoin(url, b["url"])})
url = urljoin(url, data["next"]) if data["next"] else None
with open("ratings.csv", "w", newline="", encoding="utf-8") as f:
w = csv.DictWriter(f, fieldnames=["title", "stars", "url"])
w.writeheader()
w.writerows(rows)
rated = [r["stars"] for r in rows if r["stars"]]
print(f"{len(rows)} books, average {sum(rated) / len(rated):.2f} stars")The Mystery category spans two pages. Our run fetched both (20 books, then 12, after which next was null) and printed 32 books, average 2.94 stars. The split was 7 one-star, 5 two-star, 8 three-star, 7 four-star and 5 five-star books.
For other sites, only stars() changes. For data-rating="4" read that attribute and use float(); for aria-label="4.5 out of 5 stars" take the first number with a regular expression; for a width percentage, divide by 20.
Review blocks: author, rating, date, text
Books to Scrape has ratings but no written reviews, so this part uses a made-up shop, shop.example, with typical markup. A review list is the same pattern as a product list: one container per review, and the fields inside it.
curl https://scrape.land/v1/extract \
-H "X-Api-Key: YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://shop.example/product/42/reviews?page=1",
"fields": {
"reviews": {
"css": ".review",
"fields": {
"author": ".review-author",
"rating": {"css": "[data-rating]", "attr": "data-rating"},
"date": "time@datetime",
"title": ".review-title",
"text": ".review-body"
}
},
"next": "a[rel=next]@href"
}}'The response has the shape {"data": {"reviews": [{"author": …, "rating": …, "date": …, "title": …, "text": …}, …], "next": …}}. Because each review is read from its own container, a review without a title gets "title": null in its own row and the rest stay aligned. Read the date from the datetime attribute where there is one: the visible text is often "2 weeks ago", which is useless a month later.
Reviews are usually paginated separately from the product page, either with a next link, as above, or a page number in the URL. Loop until next is null or a page returns no reviews, exactly like the ratings script, and keep a page limit in the loop as a safety net. If the review list is empty but you see reviews in your browser, the shop loads them with JavaScript; add "render": true, or look for the review API the page calls and fetch that instead.
The summary rating from JSON-LD
Many shops publish a schema.org Product in JSON-LD for search engines, and it often carries an aggregateRating with ratingValue and reviewCount, sometimes with a few review objects too. That is the exact average and count, with no rounding to half stars. POST /v1/fetch with "metadata": true returns the page's JSON-LD already parsed, under metadata.jsonld:
import requests
r = requests.post(
"https://scrape.land/v1/fetch",
headers={"X-Api-Key": "YOUR_KEY"},
json={"url": "https://shop.example/product/42", "format": "text", "metadata": True},
timeout=60,
)
r.raise_for_status()
for block in r.json().get("metadata", {}).get("jsonld", []):
for node in block.get("@graph", [block]):
agg = node.get("aggregateRating")
if agg:
print(node.get("name"), agg.get("ratingValue"), "from", agg.get("reviewCount"), "reviews")Not every shop has it. We checked a Books to Scrape product page this way: its metadata had a title, description and language, but no jsonld at all, so on that site the star class is the only source. Fetch one product page with metadata before deciding which approach to build on.
Using review data responsibly
Reviews are written by people, and many include names. Aggregate and analyse them, but think twice before republishing review text or reviewer names, check the site's terms, and follow data protection law where you and the reviewers are.
What it costs
Each page is 1 request unit, and you only pay for responses that land. The ratings run above was 2 units for 32 books. A product with 500 reviews at 20 per page is 25 units. A rendered page is 5 units, or 1 with "block_resources": true. See pricing.
Next steps
For the same group-and-paginate pattern on product listings, read scraping a paginated product catalog. For more on structured data, see page metadata and JSON-LD. Every field form is in the docs, and you can sign up free to run the script above.