arlen/benchOpen benchmarks for agentic consumers
INDEPENDENT · CC-BY-4.0
UPDATED 14 JUN 2026 · BERKELEY, CA
§ 05 · Instrument · Prototype (illustrative)

Agent Harness: Autonomous API Onboarding — Prototype

Homepage → API key → verified answers, zero humans. Identical task across Claude Code, Codex, Gemini CLI; n=5 trials per cell, clean container each, versions pinned per snapshot.

Prototype — illustrative data. No verified vendor run for this leaderboard yet; every figure is a placeholder and must not be cited as a measured result. The first verified leaderboard is web extraction.
Trials this snapshot
planned: 3 frameworks × 6 vendors × 5 trials; not yet run
Agent-ready vendors
pending first run
Fastest onboarding
pending first run
Failure taxonomy
auth_wall · schema · timeout · rate_limit
categories to be measured (no counts yet)
§ 05.1

Completion by Framework × Vendor

pending first run

No verified run yet — every cell reads — pending. When the harness runs, this section will report, for each Framework × Vendor cell (Claude Code / Codex / Gemini CLI × the scored vendors), the completion count out of 5 trials, the median time-to-first-200, the outcome or failure class (from the published taxonomy: auth_wall · schema_confusion · timeout · rate_limit), and the last-run UTC timestamp. No onboarding times, completion counts, or failure rates are published until a scored run lands — and this instrument page is noindex until then.

Vendors whose ToS prohibit automated signup are excluded from the harness and scored on the static rubric only; exclusions and reasons are published.