Epoch capability index
ECI scoreToo close to call: intervals overlap #1.
Fully AI-automated. No human reviews items before they're published. Every item links to its source. AI checks can miss errors. Report corrections.
Three rankings. Different questions. See who's ahead, how close it is, and who has the compute.
Right now: Anthropic models hold the top score in two of three rankings; no decisive margin is established.
Highest published scores, with uncertainty in view. Scales are not comparable.
Too close to call: intervals overlap #1.
No CI published. Size of the lead is uncertain.
Too close to call: intervals overlap #1.
Three different questions. Three different scales. A top score is a point estimate, not a universal winner.
90% intervals; bars span 159 to 173 within this index.
What it measuresGeneral capability inferred from a collection of benchmark results.
Source date: not published; scoring dates not provided.
Checked .
No exact row in this published snapshot. Evaluation status is unknown.
No published confidence intervals; the size of the lead is uncertain.
What it measuresTask performance across equally weighted categories, including coding and reasoning.
Source date: 2026-09-22 (file modified).
Question edition: 2026-06-25. Scores may update between editions.
Checked .
No exact row in this published snapshot. Evaluation status is unknown.
95% intervals; bars span 1489 to 1511 within this index.
What it measuresHuman preference for text responses, with style controls applied.
Source date: 2026-09-13 (leaderboard published).
Checked .
No exact row in this published snapshot. Evaluation status is unknown.
Owns: company chip fleets. Uses: lab capacity, including rentals. Epoch AI estimates, on one linear H100-equivalent scale.
Different dates and scopes. Historical estimates, not live inventory. Never add or subtract ownership and use.
→ documented use, amount unknown
Thin whiskers show usage p5 to p95. Missing estimates are not zero. Ownership source / Usage source / CC BY 4.0.
Owners: Estimates published Apr 22, 2026 · coverage through Dec 31, 2025 to Mar 31, 2026 (varies by company). Uses: Estimates published Sep 9, 2026 · coverage through Dec 31, 2025. Full compute scoreboard and event ledger →
Each lab's best published score in each index. These may be different models, with different release dates.
| Lab | Epochnot published | LiveBench2026-09-22 | LMArena2026-09-13 |
|---|---|---|---|
| 155.4DeepSeek V4 Pro 0813Published label: DeepSeek V4 Pro 0813 | 81.1DeepSeek V4.1 Flashmax, LiveBench labelPublished label: deepseek-v4.1-flash-max | 1463.4DeepSeek V4 Prohigh, 20260813 snapshot, LMArena labelPublished label: deepseek-v4-pro-high-20260813 | |
| 157.7Kimi K3Published label: Kimi K3 | 79.2Kimi K3LiveBench labelPublished label: kimi-k3 | 1484.8Kimi K3max, LMArena labelPublished label: kimi-k3-max | |
| 156.7Qwen 3.8 MaxPublished label: Qwen 3.8 Max | 78.5Qwen 3.8 MaxLiveBench labelPublished label: qwen3.8-max | 1480.6Qwen 3.8 MaxLMArena labelPublished label: qwen3.8-max | |
| 155.6GLM-5.3Published label: GLM-5.3 | 76.1GLM-5.3LiveBench labelPublished label: glm-5.3 | 1483.0GLM-5.3max, LMArena labelPublished label: glm-5.3-max | |
| 147.0MiniMax-M3Published label: MiniMax-M3 | 67.3MiniMax-M3LiveBench labelPublished label: minimax-m3 | 1441.3MiniMax-M3LMArena labelPublished label: minimax-m3 |
Overlapping intervals do not prove equality or replace a pairwise significance test. A count of first-place results is descriptive, not a combined intelligence score. The indices may share benchmarks or models.
Flagship coverage reviewed 2026-09-24. LiveBench score adaptations use CC BY-SA 4.0. Calculation and coverage notes
Epoch capability index: All scored model rows; original display labels; score order.
LiveBench: All CSV configurations; mean of seven category means; rounded to two decimals as on source. Adapted score material licensed CC BY-SA 4.0.
LMArena human preference: Published text_style_control latest, overall category only; original model labels and published rating intervals; score order.
Display names use reviewed, exact-label mappings; effort and snapshot variants remain visible. Hover or focus a name for its exact published label, also preserved in the complete table below. Generic model rows are never reassigned to a specific newer snapshot. The first five model point estimates are shown in each index; ties keep the same displayed rank. No scales are averaged. Select a lab to highlight the evidence already shown; this does not change ranking or content.
| Rank | Exact published model label | Lab | Score | Published CI |
|---|---|---|---|---|
| 1 | GPT-6 Astra | OpenAI | 166.6 | 163 to 172.03 |
| 2 | Claude Fable 5.1 | Anthropic | 165 | 161.64 to 169.61 |
| 3 | Claude Fable 5 | Anthropic | 163.6 | 160.56 to 167.57 |
| 4 | Claude Opus 5 | Anthropic | 162.67 | 159.96 to 166.52 |
| 5 | GPT-5.5 Pro | OpenAI | 162.45 | 159.23 to 166.64 |
| 6 | GPT-5.6 Sol | OpenAI | 161.99 | 159.64 to 165.92 |
| 11 | Gemini 3.7 Flash | 157.72 | 155.64 to 160.59 | |
| 12 | Kimi K3 | Moonshot | 157.68 | 155.12 to 160.65 |
| 15 | Muse Spark 1.3 | Meta | 156.89 | 154.66 to 159.44 |
| 17 | Qwen 3.8 Max | Alibaba | 156.69 | 154.56 to 158.97 |
| 18 | Grok 4.6 | SpaceXAI | 156.48 | 154.62 to 158.98 |
| 22 | GLM-5.3 | Zhipu / Z.ai | 155.56 | 153.6 to 157.99 |
| 24 | DeepSeek V4 Pro 0813 | DeepSeek | 155.39 | 153.7 to 157.31 |
| 62 | MiniMax-M3 | MiniMax | 147 | 142.93 to 149.66 |
| 179 | Amazon Nova Pro | Amazon | 123.75 | 109.68 to 126.76 |
| 187 | phi-3-small 7.4B | Microsoft | 121.79 | 113.63 to 124.99 |
| Rank | Exact published model label | Lab | Score | Published CI |
|---|---|---|---|---|
| 1 | claude-fable-5-1-max-effort | Anthropic | 83.41 | No CI published |
| 2 | claude-opus-5-5-max-effort | Anthropic | 83.22 | No CI published |
| 3 | claude-fable-5-max-effort | Anthropic | 82.97 | No CI published |
| 4 | gpt-6-astra-max | OpenAI | 82.16 | No CI published |
| 5 | claude-opus-5-5-xhigh-effort | Anthropic | 82.06 | No CI published |
| 6 | muse-spark-1.3-xhigh | Meta | 81.59 | No CI published |
| 7 | deepseek-v4.1-flash-max | DeepSeek | 81.11 | No CI published |
| 13 | kimi-k3 | Moonshot | 79.19 | No CI published |
| 14 | gemini-3.7-flash-high | 78.83 | No CI published | |
| 15 | qwen3.8-max | Alibaba | 78.46 | No CI published |
| 16 | grok-4.6 | SpaceXAI | 78.04 | No CI published |
| 29 | glm-5.3 | Zhipu / Z.ai | 76.14 | No CI published |
| 58 | minimax-m3 | MiniMax | 67.26 | No CI published |
| Rank | Exact published model label | Lab | Score | Published CI |
|---|---|---|---|---|
| 1 | claude-fable-5 | Anthropic | 1505.6827180827381 | 1500.9186317406866 to 1510.4468044247901 |
| 2 | claude-opus-4-6-high | Anthropic | 1504.5596374157553 | 1501.0420504392455 to 1508.0772243922652 |
| 3 | claude-opus-4-7-high | Anthropic | 1501.7492535919955 | 1497.8703617011465 to 1505.6281454828445 |
| 4 | muse-spark-1.2 (xHigh) | Meta | 1499.5893180176977 | 1489.081434731639 to 1510.0972013037558 |
| 5 | claude-fable-5.1-max | Anthropic | 1498.4730733605907 | 1490.3058189422723 to 1506.6403277789095 |
| 8 | muse-spark-1.3-max | Meta | 1493.1148813511777 | 1484.2701926818113 to 1501.959570020544 |
| 9 | gemini-3.8-flash-high | 1493.0082765047282 | 1484.4247213078988 to 1501.591831701557 | |
| 17 | kimi-k3-max | Moonshot | 1484.7659884633672 | 1479.540297028585 to 1489.9916798981494 |
| 18 | gpt-5.6-sol-xhigh | OpenAI | 1483.474347288669 | 1478.5808753141946 to 1488.3678192631437 |
| 19 | glm-5.3-max | Zhipu / Z.ai | 1483.0179355440732 | 1476.5775189501373 to 1489.458352138009 |
| 22 | qwen3.8-max | Alibaba | 1480.5610977215138 | 1474.770506095319 to 1486.351689347709 |
| 30 | grok-4.20-beta1 | SpaceXAI | 1474.5048857344736 | 1469.853926875293 to 1479.1558445936544 |
| 50 | deepseek-v4-pro-high-20260813 | DeepSeek | 1463.3614595040572 | 1456.5245827909412 to 1470.1983362171736 |
| 84 | minimax-m3 | MiniMax | 1441.2503431504656 | 1437.0385594969212 to 1445.46212680401 |
| 108 | amazon-nova-experimental-chat-26-02-10 | amazon | 1425.916398410703 | 1416.0436650971678 to 1435.7891317242384 |
| 306 | phi-4 | microsoft | 1256.1314601192953 | 1251.5225426068844 to 1260.740377631706 |
Find this useful? Support the project behind it.