Solutions · AI & Agent Systems
Agent Benchmark Matrix
A provenance-first matrix of agent benchmarks, showing exactly which model and harness combinations were evaluated, under what protocol, and where evidence is still missing. No ranking across unlike runs, no invented composite score.
Looking for the quality-versus-cost view instead? See the LLM Pareto Frontier. This page answers a different question: what evidence exists that a model or system can actually do the work.
Snapshot generated 13 Sept 2026 · 23 benchmarks · 22 models · 111 reported results · 3 names queued for triage.
Benchmarks tracked
19 P0 · 23 total
Models in registry
22
Result records
111 reported
Evidence split
12 official · 99 vendor
| Model | |||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
12 rows × 21 columns. A dash means no result has been ingested for that cell; nothing is ever inferred from a neighbouring model or benchmark. Rank (system) is computed only inside a comparable group.
Coverage
Coverage answers “which of my tracked models have actually been evaluated broadly?”, and where the public record is thin. It is deliberately not a quality score.
| Model | Benchmarks with a result | Official | Vendor-reported |
|---|---|---|---|
| GPT-5.6 Sol | 14 / 19 | 0 | 16 |
| GPT-5.6 Terra | 5 / 19 | 0 | 5 |
| GPT-5.6 Luna | 5 / 19 | 0 | 5 |
| GPT-5.5 | 12 / 19 | 2 | 12 |
| GLM-5.3 | 5 / 19 | 0 | 5 |
| GLM-5.3 Flash | 4 / 19 | 0 | 4 |
| GLM-5.2 | 5 / 19 | 2 | 4 |
| DeepSeek V4.1 Flash | 4 / 19 | 0 | 4 |
| DeepSeek V4 Pro | 4 / 19 | 0 | 4 |
| Kimi K3 | 13 / 19 | 0 | 13 |
| Kimi K2.6 | 0 / 19 | 0 | 0 |
| Qwen3.8 Max | 3 / 19 | 2 | 2 |
| Qwen3.8-27B | 1 / 19 | 0 | 1 |
| Qwen3.8-Flash-Next | 0 / 19 | 0 | 0 |
| Qwen3.7-Plus | 0 / 19 | 0 | 0 |
Workload axes covered
- Cross-application work12
- Tool / API / MCP8
- Browser / web6
- Enterprise SaaS6
- Deep research / browsing5
- Policy / multi-turn3
- Desktop / CUA2
- Skills / harness use1
- Terminal / coding1
Known audit caveats
- OSWorld 2.0 — 18 major, 25 minor findings (v2026.08.08). Audit
Maintenance and source diagnostics
Source freshness
- OSWorld 2.0 project page · manual · checked 2026-09-13 — Vendor/benchmark tables read manually; no machine-readable result file published yet.
- OpenAI GPT-5.6 release benchmarks · manual · checked 2026-09-13 — Vendor-official computer-use, tool-use and research tables with footnoted harness settings.
- Z.ai GLM-5.3-Flash model card · manual · checked 2026-09-13 — Vendor-official benchmark table; harness annotations vary per benchmark.
- DeepSeek-V4.1-Flash model card · manual · checked 2026-09-13 — Vendor-official table including a DeepSeek-harness comparison row.
- Kimi K3 model card · manual · checked 2026-09-13 — Vendor-official table with explicit per-benchmark harness footnotes (Kimi Code / Claude Code / Codex).
- Terminal-Bench leaderboard · manual · checked 2026-09-13 — Leaderboard submissions pin dataset, agent, model and reasoning effort.
- WebArena-Verified repo · pending · checked 2026-09-13 — Result artefacts published per release; automated adapter not yet enabled.
- τ³/τ² submissions tree · pending · checked 2026-09-13 — Structured submission.json per run; adapter planned.
- TheAgentCompany experiments · pending · checked 2026-09-13 — Versioned per-run task-level JSON; adapter planned.
- AppWorld raw leaderboard JSON · pending · checked 2026-09-13 — Canonical machine-readable source; adapter planned.
- Artificial Analysis evaluations · manual · checked 2026-09-13 — Independent evaluator; treated as a separate protocol wherever it reimplements a benchmark.
- Parsewave OSWorld 2.0 audit · manual · checked 2026-09-13 — Independent audit of release v2026.08.08: 18 major + 25 minor findings across 43 tasks.
Unresolved benchmark names (2)
- GAIA-2 — Referenced but no canonical benchmark identity established yet; queue for explicit triage.
- "τ-bench" (unversioned) — Ambiguous generation (τ vs τ² vs τ³); requires explicit version before it can be ingested.
Unresolved model names (1)
- GPT-5.6 Sol Ultra — Same model evaluated under a higher reasoning/agentic config; stored as a config on gpt-5.6-sol rather than a separate model.
Candidate benchmark registry
39 benchmarks recorded. 6 p0 · 17 p1 · 11 p2 · 3 rejected · 2 superseded.
- Mind2Web (original) — superseded: Static offline tasks; largely replaced by Online-Mind2Web and WebArena-Verified.
- MiniWoB++ — superseded: Toy-scale; retained only as a historical sanity check.
- CoWorkBench (vendor-internal) — rejected: Vendor-internal benchmark with no public tasks/protocol; cannot be a canonical comparable source.
- JobBench (vendor-internal) — rejected: Vendor-internal; may be stored as vendor-reported evidence but not a public benchmark.
- Vision2Web (vendor-internal) — rejected: Vendor-internal; no public tasks or evaluation protocol.
Provenance key
Official benchmark-official ·Paper paper ·Independent independent ·Vendor vendor-reported
How to read this
A benchmark score is never just “model → number”. It is a combination of benchmark, benchmark version, subset, metric, model revision, harness, tool and observation mode, and reasoning or step budget. Two runs of the same model under different harnesses stay separate records, and are only ranked together inside a strictly comparable group.
Missing cells stay missing. A blank is not a zero and is never filled from a neighbouring model, a family member, a vendor claim or a different benchmark. Some blanks mean “not publicly reported”; others mean “not ingested yet”. Both are stated rather than hidden.
There is deliberately no cross-benchmark “agent intelligence” score. Percentages from different benchmarks use different tasks, denominators and protocols and are not additive. Per-benchmark ranks, percentiles and coverage counts are shown instead. Provenance is labelled on every cell: benchmark-official data is the strongest class, a vendor-reported score is marked as such, and independent reimplementations are kept as their own protocols.