Skip to content
All posts

How to rank the most relevant links on a page with AI

AI extraction4 min read

A hub page (a homepage, a category index, a documentation table of contents) can link to dozens or hundreds of pages, and usually only a few of them matter for what you are doing. Instead of fetching all of them, POST /v1/rank reads the page's links and has a language model order them by relevance to a goal you describe, each with a score and a short reason. You fetch the top few and skip the rest.

The request

Give it a url and a query in plain English. top_k caps how many ranked links come back (at most 50).

curl
curl https://scrape.land/v1/rank \
  -H "X-Api-Key: YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://books.toscrape.com/",
       "query": "books and categories about travel or history",
       "top_k": 5}'

We ran it against the books.toscrape.com homepage, which has 73 links: 50 categories, 20 books, pagination and navigation.

Terminal showing the /v1/rank request and a response ranking five links: the Travel and History categories with score 1, Historical 0.95, Historical Fiction 0.9, and the travel book It's Only the Himalayas 0.85, each with a reason
The real response: total_links is 73 and the five links that match the goal come first, including a travel book that is not in a travel-named category.

The fifth result is the interesting one. "It's Only the Himalayas" is a book link, not a category, and nothing in its URL says "travel". The model ranked it because its title is about a journey. That is the kind of match a keyword filter on URLs would miss.

The response

A focused crawl in Python

Rank once, then fetch only the links that scored well:

Python
import requests

API = "https://scrape.land"
HEADERS = {"X-Api-Key": "YOUR_KEY"}

ranked = requests.post(f"{API}/v1/rank", headers=HEADERS, timeout=120, json={
    "url": "https://books.toscrape.com/",
    "query": "books and categories about travel or history",
    "top_k": 10,
}).json()

if "ranking_error" in ranked:
    print("ranking failed, links are in page order:", ranked["ranking_error"])

good = [l["url"] for l in ranked["ranked"] if l.get("score", 0) >= 0.8]
pages = requests.post(f"{API}/v1/batch", headers=HEADERS, timeout=150, json={
    "urls": good[:20],
    "format": "markdown",
}).json()["results"]

for p in pages:
    print(p["url"], len(p.get("markdown", "")), "chars")

Good uses

Ranking is stateless: it ranks the links a page already has and never fetches them. What happens next is up to your code. If you just need all links, "links": true on a fetch is cheaper; see extract every link on a page.

All the shared request options apply to the page being ranked: country, render for pages that build their links with JavaScript, and so on.

What it costs

A ranking bills the fetch (1 unit) plus the AI tier: fast (the default) adds from 2 units, smart from 4, max from 25, more for a very large page. The call above was 3 units. If the ranking step fails, only the fetch is billed. AI features, including /v1/rank, are available from the Scale plan ($499 for 3.4M requests a month) up. See pricing.

Next steps

The rank endpoint is documented in the docs. To extract typed fields from the pages you picked, see AI schema extraction. Create an account to get started.

Start free with 1,000 requests Read the docs