Leaderboards
Ranked tables for every web-access capability we benchmark — composite score, quality metric, latency p50 and cost per success for each provider, with run dates.
What this board tests: a natural-language goal handed to an autonomous browser agent.
What this board tests: structured fields pulled from a page against an ai/natural-language schema.
What this board tests: declarative in-page interactions (click, fill, scroll) in one call.
What this board tests: change detection / monitoring on a url.
What this board tests: pdf/document to text, with ocr where the provider supports it.
What this board tests: purpose-built scrapers for hard, high-value domains (e-commerce, social).
What this board tests: structured fields pulled from a page via css/xpath rules.
What this board tests: rendered image of a url.
What this board tests: google serp verticals (web, news, images, places, scholar) as structured json.
What this board tests: enumerate and fetch many pages across a site.
What this board tests: a question in, a synthesized answer with cited sources out.
What this board tests: fetch one url as clean content/markdown, anti-bot handled.
What this board tests: web search: a query in, ranked result links out.
How to read these numbers
One shared corpus per capability, identical for every provider; composites blend measured quality, latency p50, cost per successful call and error rate. Nobody gets dropped for scoring badly, and no score is ever adjusted by hand — the point of the exercise is that you can trust the order. Full detail on corpora, metrics and grading: the methodology page.