How to rank the most relevant links on a page with AI
A hub page (a homepage, a category index, a documentation table of contents) can link to dozens or hundreds of pages, and usually only a few of them matter for what you are doing. Instead of fetching all of them, POST /v1/rank reads the page's links and has a language model order them by relevance to a goal you describe, each with a score and a short reason. You fetch the top few and skip the rest.
The request
Give it a url and a query in plain English. top_k caps how many ranked links come back (at most 50).
curl https://scrape.land/v1/rank \
-H "X-Api-Key: YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://books.toscrape.com/",
"query": "books and categories about travel or history",
"top_k": 5}'We ran it against the books.toscrape.com homepage, which has 73 links: 50 categories, 20 books, pagination and navigation.

total_links is 73 and the five links that match the goal come first, including a travel book that is not in a travel-named category.The fifth result is the interesting one. "It's Only the Himalayas" is a book link, not a category, and nothing in its URL says "travel". The model ranked it because its title is about a journey. That is the kind of match a keyword filter on URLs would miss.
The response
ranked: the links in order, each with the absoluteurl, the link's anchortext, ascorefrom 0 to 1 and a one-linereason.total_links: how many links the page had;ranked_count: how many came back.model: the AI tier that answered. If you ask for a tier that is unavailable, the API answers on another one and says so withrequested_modelnext tomodel.ranking_error: only present if the AI step failed. You still get a200with the page's links in document order, so your code keeps working.
A focused crawl in Python
Rank once, then fetch only the links that scored well:
import requests
API = "https://scrape.land"
HEADERS = {"X-Api-Key": "YOUR_KEY"}
ranked = requests.post(f"{API}/v1/rank", headers=HEADERS, timeout=120, json={
"url": "https://books.toscrape.com/",
"query": "books and categories about travel or history",
"top_k": 10,
}).json()
if "ranking_error" in ranked:
print("ranking failed, links are in page order:", ranked["ranking_error"])
good = [l["url"] for l in ranked["ranked"] if l.get("score", 0) >= 0.8]
pages = requests.post(f"{API}/v1/batch", headers=HEADERS, timeout=150, json={
"urls": good[:20],
"format": "markdown",
}).json()["results"]
for p in pages:
print(p["url"], len(p.get("markdown", "")), "chars")Good uses
- Research agents. Given a company homepage, find the pricing, careers or contact page without hardcoding paths.
- Documentation and knowledge bases. Pick the pages about one feature out of a large table of contents before you fetch them as Markdown.
- Lead enrichment. From a business's homepage, rank for "contact page, team page, about us" and read only those. See directory page to lead list.
Ranking is stateless: it ranks the links a page already has and never fetches them. What happens next is up to your code. If you just need all links, "links": true on a fetch is cheaper; see extract every link on a page.
All the shared request options apply to the page being ranked: country, render for pages that build their links with JavaScript, and so on.
What it costs
A ranking bills the fetch (1 unit) plus the AI tier: fast (the default) adds from 2 units, smart from 4, max from 25, more for a very large page. The call above was 3 units. If the ranking step fails, only the fetch is billed. AI features, including /v1/rank, are available from the Scale plan ($499 for 3.4M requests a month) up. See pricing.
Next steps
The rank endpoint is documented in the docs. To extract typed fields from the pages you picked, see AI schema extraction. Create an account to get started.