Leaderboards
Ranked tables for every web-access capability we benchmark — composite score, quality metric, latency p50 and the track-specific cost metric for each provider, with run dates.
What this board tests: a natural-language goal handed to an autonomous browser agent.
What this board tests: structured fields pulled from a page against an ai/natural-language schema.
What this board tests: declarative in-page interactions (click, fill, scroll) in one call.
What this board tests: change detection / monitoring on a url.
What this board tests: pdf/document to text, with ocr where the provider supports it.
What this board tests: purpose-built scrapers for hard, high-value domains (e-commerce, social).
What this board tests: structured fields pulled from a page via css/xpath rules.
What this board tests: rendered image of a url.
What this board tests: google serp verticals (web, news, images, places, scholar) as structured json.
What this board tests: enumerate and fetch many pages across a site.
What this board tests: a question in, a synthesized answer with cited sources out.
What this board tests: fetch one url as clean content/markdown, anti-bot handled.
What this board tests: web search: a query in, ranked result links out.
The run behind these boards
Every table above comes from the benchmark run of 2026-09-14: 67 scorecards across 13 capabilities, covering 22 of the 34 providers in the catalog. Other catalog entries either fall outside these benchmark capabilities or carry an explicit unscored status on their provider page. Per-provider numbers and raw metrics: the catalog and evals.json.
Not scored in this run
Every unscored pair recorded for this run is listed explicitly. Source-only benchmark pairs are run provenance, not NativePort catalog entries:
- Brave Search
—
brave:answer· Answer: not scored (unscored_in_run), measured 2026-09-14. - Oxylabs
—
oxylabs:crawl· Crawl: not scored (unscored_in_run), measured 2026-09-14. - Oxylabs
—
oxylabs:act_agent· Act · NL-agent: not scored (unscored_in_run), measured 2026-09-14. youtube-data-api(source-only benchmark pair; not a catalog entry) —youtube-data-api:scrape_domain· Scrape-domain: not scored (unscored_in_run), measured 2026-09-14.google-books-api(source-only benchmark pair; not a catalog entry) —google-books-api:scrape_domain· Scrape-domain: not scored (unscored_in_run), measured 2026-09-14.
How to read these numbers
One shared corpus per capability, identical for every provider; composites blend measured quality, latency p50, the track-specific cost metric and error rate. Nobody gets dropped for scoring badly, and no score is ever adjusted by hand — the point of the exercise is that you can trust the order. Full detail on corpora, metrics and grading: the methodology page.