Solutions · AI & Agent Systems

Agent Benchmark Matrix

A provenance-first matrix of agent benchmarks, showing exactly which model and harness combinations were evaluated, under what protocol, and where evidence is still missing. No ranking across unlike runs, no invented composite score.

Looking for the quality-versus-cost view instead? See the LLM Pareto Frontier. This page answers a different question: what evidence exists that a model or system can actually do the work.

Snapshot generated 13 Sept 2026 · 23 benchmarks · 22 models · 111 reported results · 3 names queued for triage.

Benchmarks tracked

19 P0 · 23 total

Models in registry

22

Result records

111 reported

Evidence split

12 official · 99 vendor

View
Comparability
Evidence
Cohort
Model or system coverage across tracked agent benchmarks. Each cell shows the best credible score, its provenance class and its rank within its comparable group.
Model

12 rows × 21 columns. A dash means no result has been ingested for that cell; nothing is ever inferred from a neighbouring model or benchmark. Rank (system) is computed only inside a comparable group.

Coverage

Coverage answers “which of my tracked models have actually been evaluated broadly?”, and where the public record is thin. It is deliberately not a quality score.

ModelBenchmarks with a resultOfficialVendor-reported
GPT-5.6 Sol14 / 19016
GPT-5.6 Terra5 / 1905
GPT-5.6 Luna5 / 1905
GPT-5.512 / 19212
GLM-5.35 / 1905
GLM-5.3 Flash4 / 1904
GLM-5.25 / 1924
DeepSeek V4.1 Flash4 / 1904
DeepSeek V4 Pro4 / 1904
Kimi K313 / 19013
Kimi K2.60 / 1900
Qwen3.8 Max3 / 1922
Qwen3.8-27B1 / 1901
Qwen3.8-Flash-Next0 / 1900
Qwen3.7-Plus0 / 1900

Workload axes covered

  • Cross-application work12
  • Tool / API / MCP8
  • Browser / web6
  • Enterprise SaaS6
  • Deep research / browsing5
  • Policy / multi-turn3
  • Desktop / CUA2
  • Skills / harness use1
  • Terminal / coding1

Known audit caveats

  • OSWorld 2.018 major, 25 minor findings (v2026.08.08). Audit
Maintenance and source diagnostics

Source freshness

  • OSWorld 2.0 project page · manual · checked 2026-09-13 — Vendor/benchmark tables read manually; no machine-readable result file published yet.
  • OpenAI GPT-5.6 release benchmarks · manual · checked 2026-09-13 — Vendor-official computer-use, tool-use and research tables with footnoted harness settings.
  • Z.ai GLM-5.3-Flash model card · manual · checked 2026-09-13 — Vendor-official benchmark table; harness annotations vary per benchmark.
  • DeepSeek-V4.1-Flash model card · manual · checked 2026-09-13 — Vendor-official table including a DeepSeek-harness comparison row.
  • Kimi K3 model card · manual · checked 2026-09-13 — Vendor-official table with explicit per-benchmark harness footnotes (Kimi Code / Claude Code / Codex).
  • Terminal-Bench leaderboard · manual · checked 2026-09-13 — Leaderboard submissions pin dataset, agent, model and reasoning effort.
  • WebArena-Verified repo · pending · checked 2026-09-13 — Result artefacts published per release; automated adapter not yet enabled.
  • τ³/τ² submissions tree · pending · checked 2026-09-13 — Structured submission.json per run; adapter planned.
  • TheAgentCompany experiments · pending · checked 2026-09-13 — Versioned per-run task-level JSON; adapter planned.
  • AppWorld raw leaderboard JSON · pending · checked 2026-09-13 — Canonical machine-readable source; adapter planned.
  • Artificial Analysis evaluations · manual · checked 2026-09-13 — Independent evaluator; treated as a separate protocol wherever it reimplements a benchmark.
  • Parsewave OSWorld 2.0 audit · manual · checked 2026-09-13 — Independent audit of release v2026.08.08: 18 major + 25 minor findings across 43 tasks.

Unresolved benchmark names (2)

  • GAIA-2Referenced but no canonical benchmark identity established yet; queue for explicit triage.
  • "τ-bench" (unversioned)Ambiguous generation (τ vs τ² vs τ³); requires explicit version before it can be ingested.

Unresolved model names (1)

  • GPT-5.6 Sol UltraSame model evaluated under a higher reasoning/agentic config; stored as a config on gpt-5.6-sol rather than a separate model.

Candidate benchmark registry

39 benchmarks recorded. 6 p0 · 17 p1 · 11 p2 · 3 rejected · 2 superseded.

  • Mind2Web (original)superseded: Static offline tasks; largely replaced by Online-Mind2Web and WebArena-Verified.
  • MiniWoB++superseded: Toy-scale; retained only as a historical sanity check.
  • CoWorkBench (vendor-internal)rejected: Vendor-internal benchmark with no public tasks/protocol; cannot be a canonical comparable source.
  • JobBench (vendor-internal)rejected: Vendor-internal; may be stored as vendor-reported evidence but not a public benchmark.
  • Vision2Web (vendor-internal)rejected: Vendor-internal; no public tasks or evaluation protocol.

Provenance key

Official benchmark-official ·Paper paper ·Independent independent ·Vendor vendor-reported

How to read this

A benchmark score is never just “model → number”. It is a combination of benchmark, benchmark version, subset, metric, model revision, harness, tool and observation mode, and reasoning or step budget. Two runs of the same model under different harnesses stay separate records, and are only ranked together inside a strictly comparable group.

Missing cells stay missing. A blank is not a zero and is never filled from a neighbouring model, a family member, a vendor claim or a different benchmark. Some blanks mean “not publicly reported”; others mean “not ingested yet”. Both are stated rather than hidden.

There is deliberately no cross-benchmark “agent intelligence” score. Percentages from different benchmarks use different tasks, denominators and protocols and are not additive. Per-benchmark ranks, percentiles and coverage counts are shown instead. Provenance is labelled on every cell: benchmark-official data is the strongest class, a vendor-reported score is marked as such, and independent reimplementations are kept as their own protocols.