SUBSTRATE WIRE

Fully AI-automated. No human reviews items before they're published. Every item links to its source. AI checks can miss errors. Report corrections.

The AI race

Three rankings. Different questions. See who's ahead, how close it is, and who has the compute.

Right now: Anthropic models hold the top score in two of three rankings; no decisive margin is established.

The podium

Highest published scores, with uncertainty in view. Scales are not comparable.

Epoch capability index

ECI score
166.6Highest point estimate
165.0
163.6
162.7
162.4
162.0

Too close to call: intervals overlap #1.

Measures: General capability inferred from a collection of benchmark results.

Source date: not published. Checked 2026-09-25 UTC.

View source

LiveBench

Overall / 100
83.4Highest point estimate

No CI published. Size of the lead is uncertain.

Measures: Task performance across equally weighted categories, including coding and reasoning.

Source date: 2026-09-22. Checked 2026-09-25 UTC.

View source

LMArena human preference

Style-controlled text rating
1505.7Highest point estimate
1504.6
1501.7
1499.6
1498.5
1493.1
1493.0

Too close to call: intervals overlap #1.

Measures: Human preference for text responses, with style controls applied.

Source date: 2026-09-13. Checked 2026-09-25 UTC.

View source

Who's smartest right now

Three different questions. Three different scales. A top score is a point estimate, not a universal winner.

Epoch capability index

ECI score
Too close to call
  1. 1
    166.6
  2. 2
    165.0
  3. 3
    163.6
  4. 4
    162.7
  5. 5
    162.4

90% intervals; bars span 159 to 173 within this index.

What it measures / Not yet scored / Best by lab

What it measuresGeneral capability inferred from a collection of benchmark results.

Source date: not published; scoring dates not provided.

Checked .

Source · CC BY 4.0

Best published model by lab

  • 166.6GPT-6 AstraPublished label: GPT-6 Astra
  • 165.0Claude Fable 5.1Published label: Claude Fable 5.1
  • 157.7Gemini 3.7 FlashPublished label: Gemini 3.7 Flash
  • 156.9Muse Spark 1.3Published label: Muse Spark 1.3
  • 156.5Grok 4.6Published label: Grok 4.6
  • 155.4DeepSeek V4 Pro 0813Published label: DeepSeek V4 Pro 0813
  • 157.7Kimi K3Published label: Kimi K3
  • 156.7Qwen 3.8 MaxPublished label: Qwen 3.8 Max
  • 155.6GLM-5.3Published label: GLM-5.3
  • 147.0MiniMax-M3Published label: MiniMax-M3
  • 121.8phi-3-small 7.4BPublished label: phi-3-small 7.4B
  • 123.8Amazon Nova ProPublished label: Amazon Nova Pro

Not yet scored 6

No exact row in this published snapshot. Evaluation status is unknown.

LiveBench

Overall / 100
  1. 1
    No CI published
    83.4
  2. 2
    No CI published
    83.2
  3. 3
    No CI published
    83.0
  4. 4
    No CI published
    82.2
  5. 5
    No CI published
    82.1

No published confidence intervals; the size of the lead is uncertain.

What it measures / Not yet scored / Best by lab

What it measuresTask performance across equally weighted categories, including coding and reasoning.

Source date: 2026-09-22 (file modified).

Question edition: 2026-06-25. Scores may update between editions.

Checked .

Source · CC BY-SA 4.0

Best published model by lab

  • 82.2GPT-6 Astramax, LiveBench labelPublished label: gpt-6-astra-max
  • 83.4Claude Fable 5.1max effort, LiveBench labelPublished label: claude-fable-5-1-max-effort
  • 78.8Gemini 3.7 Flashhigh, LiveBench labelPublished label: gemini-3.7-flash-high
  • 81.6Muse Spark 1.3xhigh, LiveBench labelPublished label: muse-spark-1.3-xhigh
  • 78.0Grok 4.6LiveBench labelPublished label: grok-4.6
  • 81.1DeepSeek V4.1 Flashmax, LiveBench labelPublished label: deepseek-v4.1-flash-max
  • 79.2Kimi K3LiveBench labelPublished label: kimi-k3
  • 78.5Qwen 3.8 MaxLiveBench labelPublished label: qwen3.8-max
  • 76.1GLM-5.3LiveBench labelPublished label: glm-5.3
  • 67.3MiniMax-M3LiveBench labelPublished label: minimax-m3

Not yet scored 4

No exact row in this published snapshot. Evaluation status is unknown.

LMArena human preference

Style-controlled text rating
Too close to call
  1. 1
    1505.7
  2. 2
    1504.6
  3. 3
    1501.7
  4. 4
    1499.6
  5. 5
    1498.5

95% intervals; bars span 1489 to 1511 within this index.

What it measures / Not yet scored / Best by lab

What it measuresHuman preference for text responses, with style controls applied.

Source date: 2026-09-13 (leaderboard published).

Checked .

Source · CC BY 4.0

Best published model by lab

  • 1483.5GPT-5.6 Solxhigh, LMArena labelPublished label: gpt-5.6-sol-xhigh
  • 1505.7Claude Fable 5LMArena labelPublished label: claude-fable-5
  • 1493.0Gemini 3.8 Flashhigh, LMArena labelPublished label: gemini-3.8-flash-high
  • 1499.6Muse Spark 1.2xHigh, LMArena labelPublished label: muse-spark-1.2 (xHigh)
  • 1474.5Grok 4.20beta1, LMArena labelPublished label: grok-4.20-beta1
  • 1463.4DeepSeek V4 Prohigh, 20260813 snapshot, LMArena labelPublished label: deepseek-v4-pro-high-20260813
  • 1484.8Kimi K3max, LMArena labelPublished label: kimi-k3-max
  • 1480.6Qwen 3.8 MaxLMArena labelPublished label: qwen3.8-max
  • 1483.0GLM-5.3max, LMArena labelPublished label: glm-5.3-max
  • 1441.3MiniMax-M3LMArena labelPublished label: minimax-m3
  • 1256.1Phi-4LMArena labelPublished label: phi-4
  • 1425.9Amazon Novaexperimental chat, 26-02-10 snapshot, LMArena labelPublished label: amazon-nova-experimental-chat-26-02-10

Not yet scored 8

No exact row in this published snapshot. Evaluation status is unknown.

Who has the power

Owns: company chip fleets. Uses: lab capacity, including rentals. Epoch AI estimates, on one linear H100-equivalent scale.

Different dates and scopes. Historical estimates, not live inventory. Never add or subtract ownership and use.

03M6M H100e
Owns
5.47M H100eMar 31, 2026 · incomplete
Uses
1.58M H100eDec 31, 2025p5 to p95: 1.01M to 2.55M H100e.
Owns
3.59M H100eDec 31, 2025 to Mar 31, 2026 · incomplete · mixed periods
Uses
not estimated
Owns
2.52M H100eDec 31, 2025 to Mar 31, 2026 · incomplete · mixed periods
Uses
not estimated
Owns
2.41M H100eDec 31, 2025 to Mar 31, 2026 · incomplete · mixed periods
Uses
996K H100eDec 31, 2025p5 to p95: 606K to 1.64M H100e.
Owns
not estimated
Uses
1.74M H100eDec 31, 2025p5 to p95: 1.25M to 2.19M H100e.
Owns
1.21M H100eDec 31, 2025 to Mar 31, 2026 · incomplete · mixed periods
Uses
not estimated
Owns
not estimated
Uses
1.19M H100eDec 31, 2025p5 to p95: 842K to 1.72M H100e.
Owns
832K H100eDec 31, 2025
Uses
not estimated
Owns
554K H100eDec 31, 2025
Uses
635K H100eDec 31, 2025p5 to p95: 587K to 787K H100e.

→ documented use, amount unknown

Thin whiskers show usage p5 to p95. Missing estimates are not zero. Ownership source / Usage source / CC BY 4.0.

Owners: Estimates published Apr 22, 2026 · coverage through Dec 31, 2025 to Mar 31, 2026 (varies by company). Uses: Estimates published Sep 9, 2026 · coverage through Dec 31, 2025. Full compute scoreboard and event ledger →

China's labs, in view

Each lab's best published score in each index. These may be different models, with different release dates.

China labs' highest model point estimates by index
LabEpochnot publishedLiveBench2026-09-22LMArena2026-09-13
155.4DeepSeek V4 Pro 0813Published label: DeepSeek V4 Pro 081381.1DeepSeek V4.1 Flashmax, LiveBench labelPublished label: deepseek-v4.1-flash-max1463.4DeepSeek V4 Prohigh, 20260813 snapshot, LMArena labelPublished label: deepseek-v4-pro-high-20260813
157.7Kimi K3Published label: Kimi K379.2Kimi K3LiveBench labelPublished label: kimi-k31484.8Kimi K3max, LMArena labelPublished label: kimi-k3-max
156.7Qwen 3.8 MaxPublished label: Qwen 3.8 Max78.5Qwen 3.8 MaxLiveBench labelPublished label: qwen3.8-max1480.6Qwen 3.8 MaxLMArena labelPublished label: qwen3.8-max
155.6GLM-5.3Published label: GLM-5.376.1GLM-5.3LiveBench labelPublished label: glm-5.31483.0GLM-5.3max, LMArena labelPublished label: glm-5.3-max
147.0MiniMax-M3Published label: MiniMax-M367.3MiniMax-M3LiveBench labelPublished label: minimax-m31441.3MiniMax-M3LMArena labelPublished label: minimax-m3
Calculation and coverage notes

Overlapping intervals do not prove equality or replace a pairwise significance test. A count of first-place results is descriptive, not a combined intelligence score. The indices may share benchmarks or models.

Flagship coverage reviewed 2026-09-24. LiveBench score adaptations use CC BY-SA 4.0. Calculation and coverage notes

Epoch capability index: All scored model rows; original display labels; score order.

LiveBench: All CSV configurations; mean of seven category means; rounded to two decimals as on source. Adapted score material licensed CC BY-SA 4.0.

LMArena human preference: Published text_style_control latest, overall category only; original model labels and published rating intervals; score order.

Display names use reviewed, exact-label mappings; effort and snapshot variants remain visible. Hover or focus a name for its exact published label, also preserved in the complete table below. Generic model rows are never reassigned to a specific newer snapshot. The first five model point estimates are shown in each index; ties keep the same displayed rank. No scales are averaged. Select a lab to highlight the evidence already shown; this does not change ranking or content.

All displayed results and exact labels
Epoch capability index: 16 displayed results
RankExact published model labelLabScorePublished CI
1GPT-6 AstraOpenAI166.6163 to 172.03
2Claude Fable 5.1Anthropic165161.64 to 169.61
3Claude Fable 5Anthropic163.6160.56 to 167.57
4Claude Opus 5Anthropic162.67159.96 to 166.52
5GPT-5.5 ProOpenAI162.45159.23 to 166.64
6GPT-5.6 SolOpenAI161.99159.64 to 165.92
11Gemini 3.7 FlashGoogle157.72155.64 to 160.59
12Kimi K3Moonshot157.68155.12 to 160.65
15Muse Spark 1.3Meta156.89154.66 to 159.44
17Qwen 3.8 MaxAlibaba156.69154.56 to 158.97
18Grok 4.6SpaceXAI156.48154.62 to 158.98
22GLM-5.3Zhipu / Z.ai155.56153.6 to 157.99
24DeepSeek V4 Pro 0813DeepSeek155.39153.7 to 157.31
62MiniMax-M3MiniMax147142.93 to 149.66
179Amazon Nova ProAmazon123.75109.68 to 126.76
187phi-3-small 7.4BMicrosoft121.79113.63 to 124.99
LiveBench: 13 displayed results
RankExact published model labelLabScorePublished CI
1claude-fable-5-1-max-effortAnthropic83.41No CI published
2claude-opus-5-5-max-effortAnthropic83.22No CI published
3claude-fable-5-max-effortAnthropic82.97No CI published
4gpt-6-astra-maxOpenAI82.16No CI published
5claude-opus-5-5-xhigh-effortAnthropic82.06No CI published
6muse-spark-1.3-xhighMeta81.59No CI published
7deepseek-v4.1-flash-maxDeepSeek81.11No CI published
13kimi-k3Moonshot79.19No CI published
14gemini-3.7-flash-highGoogle78.83No CI published
15qwen3.8-maxAlibaba78.46No CI published
16grok-4.6SpaceXAI78.04No CI published
29glm-5.3Zhipu / Z.ai76.14No CI published
58minimax-m3MiniMax67.26No CI published
LMArena human preference: 16 displayed results
RankExact published model labelLabScorePublished CI
1claude-fable-5Anthropic1505.68271808273811500.9186317406866 to 1510.4468044247901
2claude-opus-4-6-highAnthropic1504.55963741575531501.0420504392455 to 1508.0772243922652
3claude-opus-4-7-highAnthropic1501.74925359199551497.8703617011465 to 1505.6281454828445
4muse-spark-1.2 (xHigh)Meta1499.58931801769771489.081434731639 to 1510.0972013037558
5claude-fable-5.1-maxAnthropic1498.47307336059071490.3058189422723 to 1506.6403277789095
8muse-spark-1.3-maxMeta1493.11488135117771484.2701926818113 to 1501.959570020544
9gemini-3.8-flash-highGoogle1493.00827650472821484.4247213078988 to 1501.591831701557
17kimi-k3-maxMoonshot1484.76598846336721479.540297028585 to 1489.9916798981494
18gpt-5.6-sol-xhighOpenAI1483.4743472886691478.5808753141946 to 1488.3678192631437
19glm-5.3-maxZhipu / Z.ai1483.01793554407321476.5775189501373 to 1489.458352138009
22qwen3.8-maxAlibaba1480.56109772151381474.770506095319 to 1486.351689347709
30grok-4.20-beta1SpaceXAI1474.50488573447361469.853926875293 to 1479.1558445936544
50deepseek-v4-pro-high-20260813DeepSeek1463.36145950405721456.5245827909412 to 1470.1983362171736
84minimax-m3MiniMax1441.25034315046561437.0385594969212 to 1445.46212680401
108amazon-nova-experimental-chat-26-02-10amazon1425.9163984107031416.0436650971678 to 1435.7891317242384
306phi-4microsoft1256.13146011929531251.5225426068844 to 1260.740377631706

Find this useful? Support the project behind it.