NativePort

How we measure

Last updated July 23, 2026


Every number on this site — each composite, rank, latency and cost figure — comes from benchmark runs we operate ourselves. This page describes exactly what those runs do, so you can decide how much weight the numbers deserve.

One corpus per capability, identical for everyone

We benchmark web-access providers per capability: web search, SERP verticals, scraping, crawling, AI-schema extraction, rule-based extraction, sourced answers, screenshots, domain-specific scraping, declarative browser actions, autonomous browser agents, document parsing, and change watching.

Each capability has a versioned task corpus (currently v1) that is held fixed across providers: the same URLs, the same queries, the same target schemas, the same pass criteria. A scraping provider isn’t compared against a different set of pages than its competitor — ever. When a corpus version changes, the version changes for every provider at once and new run dates make the switch visible.

Corpora are built to exercise the hard parts, not the demo path: pages behind commercial anti-bot walls, scanned PDFs that require OCR, sites that block datacenter IPs, queries whose answers require synthesis across sources.

Four dimensions, measured not estimated

For every provider × capability pair we record:

  1. Quality — the capability’s own success definition: anti-bot bypass rate and markdown cleanliness for scraping, recall for search, field accuracy for extraction, coverage for crawling, task success for browser actions, text accuracy for parsing, valid-image rate for screenshots.
  2. Latency p50 — median wall-clock time per call, measured from our runner.
  3. Cost per successful call — computed from the provider’s real metered price for the calls in the run, denominated in dollars. Failures inflate this number by design: paying for errors is part of a provider’s true cost.
  4. Error rate — non-2xx responses and exceptions across the run.

Where quality requires judgment rather than string comparison — how clean is this markdown, is this extraction faithful — grading is done by an LLM panel (Claude Opus and GPT-5) working from written rubrics, so no single model’s taste decides a score.

These fold into a composite out of 10, and providers are ranked within each capability. The raw per-metric values are published alongside every composite, on the provider pages and in machine-readable form at /evals.json.

Run dates on everything

Scores go stale; ours say so. Every scorecard carries the date of the run that produced it, and pages surface the newest date backing their numbers. If a scorecard is old, that’s visible — we’d rather show a dated number than a fresh-looking guess.

What keeps the ranking honest

  • Weak scores are published. A provider that lands last on a board stays on the board. The catalog is only useful if the ordering is real.
  • Nobody can buy a position. Providers don’t pay to be listed and can’t pay to move. Most don’t know when a run happens.
  • The business model can’t lean on the scale. NativePort’s fee is 5.5% at credit top-up — flat, regardless of which provider you route to. Steering you toward an expensive provider earns us nothing, and the per-call prices themselves pass through unmarked.
  • Gated paradigms are labeled, not shoehorned. Some products can’t be scored on a request/response corpus — Browserbase, for instance, hands you a browser session to drive rather than a URL-in/content-out endpoint. We say so on its page instead of forcing a number that would mean nothing.

What we don’t claim

We measure what our runner can observe. We don’t publish uptime percentages, SLA figures, or throughput ceilings for providers — those need longitudinal infrastructure we haven’t pointed at them. Where a claim on this site isn’t backed by a run, it isn’t made.

Consuming the data