Solutions · Analytics

LLM Pareto Frontier

Deterministic quality-vs-cost frontier for major LLMs, built from official structured data: Artificial Analysis v2, OpenRouter, and LM Arena's published Hugging Face dataset. No scraping, no fuzzy matching.

One or more sources degraded; showing last-known data. AA: degraded, retry after 41264s (AA API rate limited)

Artificial Analysiserror

Fetched 2 Sept 2026, 18:14

OpenRouterok

Fetched 3 Sept 2026, 06:35

LM Arena datasetpending

Published · fetched

Mode

Model table

Links
Claude Fable 5.1Anthropic65.7$3.69$10.00$50.00OpenRouterLab
Claude Opus 5Anthropic63.1$2.34$5.00$25.00OpenRouterLab
Grok 4.6xAI60.9$0.9372$2.00$6.00OpenRouterLab
Kimi K3Moonshot AI59.7$0.8375$3.00$15.00OpenRouterLab
GLM 5.3Z.ai59.5$0.6829$1.40$4.40OpenRouterLab
Gemini 3.8 FlashGoogle58.7$0.5766$0.7500$3.75OpenRouterLab
GLM 5.3 FlashZ.ai57.5$0.0869$0.0750$0.2500OpenRouterLab
GPT-5.6 LunaOpenAI52.3$0.0487$0.2000$1.20OpenRouterLab
Hunyuan Hy3Tencent42.2$0.0357$0.1320$0.5280OpenRouterLab
Llama 4 MaverickMeta14.5$0.0346$0.2000$0.6960OpenRouterLab
Llama 4 ScoutMeta10.3$0.0106$0.1000$0.3000OpenRouterLab

Methodology

Pareto dominance: model A dominates B iff Q(A) ≥ Q(B) and C(A) ≤ C(B) and at least one inequality is strict. Models missing the selected metric are excluded from that frontier. The graph and table show only the currently selected Pareto-efficient models.

Blended OpenRouter cost = inputShare × input $/1M + outputShare × output $/1M (visible and editable above when selected).

Arena WebDev uses Bradley-Terry Arena Scores (rating ± bounds, vote counts). Arena Agent uses IPS scores (score ± CI, observation and session counts). These methodologies are kept semantically separate and are never merged into a single "Elo".