arlen/benchOpen benchmarks for agentic consumers
INDEPENDENT · CC-BY-4.0
UPDATED 14 JUN 2026 · BERKELEY, CA
Independent System of Record · CC-BY-4.0

Web Search & Extraction API Benchmarks for AI Agents

Open benchmarks for agentic consumers — the APIs AI agents buy and use mid-task.

Safe to cite today Two leaderboards are verified (pre-review): web search — Exa leads hit@5 at 80.9% (n=299) — and web extraction — Exa leads fidelity 0.74 (n=150, independently audited). Cost, freshness, KYB and the agent harness are pending or prototype — don't cite them. See the claim table ↓ · claims.json
The Verdict snapshot web_search-2026-q3 · 299 verified queries · 4 vendors

Answer: for golden-URL search accuracy, Exa leads. Caveat first: only web search and web extraction are verified — cost, freshness, KYB and the agent harness are pending or prototype.  On snapshot web_search-2026-q3 (299 public verified queries, 4 vendors; hit@1/hit@5 only this run): Exa (80.9% hit@5) is the clear leader over SerpAPI (69.2%) — their 95% confidence intervals are disjoint and McNemar p≈0.001, so the two are statistically separable at n=299 — both ahead of Brave (58.9%) and Tavily (49.4%). Cost, freshness & latency: pending — sentinel & cost instruments not yet run.

449
Rows scored & verified (150 ext + 299 search)
4
Vendors scored
2
Verified boards
1,550
Corpus curated (not all scored)
CC-BY 4.0
License
§ 00

What's verified — and safe to cite

one row per claim · status-tagged
Claim Snapshot n Vendors Status Safe to cite?
Web search — Exa leads hit@5 (80.9%), ahead of SerpAPI (69.2%)web_search-2026-q32994verified · reproducible, pre-review✓ yes
Web extraction — Exa leads main-content fidelity (0.74)web_extraction-2026-q21504verified · audited, pre-review✓ yes
Web extraction — cost-per-correct (Exa lower of the 2 priced)web_extraction-2026-q21502 pricedprovisional✗ provisional
KYB / identity — vendor accuracyfirst run July 2026✗ not yet
Freshness lag — time-to-retrievabilityprototype (illustrative)✗ no
Agent harness — autonomous onboardingprototype (illustrative)✗ no

Cite only rows marked , with their snapshot id and caveat. The machine-readable equivalent (with confidence intervals + per-claim caveats) is claims.json. "Verified" snapshots are published pre-review — vendor right-of-reply rows have not yet been sent. "Corpus curated" (1,550) counts hand-verified golden items built across all primitives; only 449 are scored on a verified leaderboard today.

§ 03

Methodology & Independence

golden truth · private holdout · published

Every leaderboard is scored against row-by-row golden truth with a private 30% holdout to deter overfitting. Snapshots are versioned and dated; claims cite the snapshot ID. No vendor can pay for placement, and vendors are entitled to a documented right of reply — corrections are published, not negotiated.

Read the full methodology →