TESTS
What each test measures.
Every published test on the site, by skill: who runs it, how many models it covers and which model scores highest.
Overall
Coding
Math
- LiveBench mathematicsLiveBench, 63 models
Top scoreClaude Opus 5.5 (max effort) 97.1%
- Mock AIME 2024–2025Epoch AI, 198 models
Top scoreClaude Fable 5 (high) 100.0%
- FrontierMath Tiers 1–3Epoch AI, 81 models
Top scoreGPT-6 Astra (max) 93.7%
- FrontierMath Tier 4Epoch AI, 63 models
Top scoreGPT-6.1 Sol (max) 100.0%
- MATH Level 5Epoch AI, 97 models
Top scoreGPT-5 (high) 98.1%
- Arena mathArena, 372 models
Top scoreGemini 4 Argon High 1530
Knowledge
Agents
- LiveBench agentic codingLiveBench, 63 models
Top scoreDeepSeek V4.1 Flash (max) 77.3%
- DeepSWEDatacurve DeepSWE, compiled by Epoch AI, 26 models
Top scoreGPT-6 Astra (xhigh) 74.1%
- CursorBenchCursor CursorBench, compiled by Epoch AI, 14 models
Top scoreClaude Opus 5.5 (max) 57.8%
- SWE-bench Verified, bash onlySWE-bench, 42 models
Top scoreClaude Opus 4.5 (high) 76.8%
Language
- LiveBench languageLiveBench, 63 models
Top scoreClaude Fable 5 (max effort) 90.7%
- LiveBench instructionsLiveBench, 63 models
Top scoreGemini 3.8 Flash (high) 81.4%
- Arena creative writingArena, 387 models
Top scoreGemini 4 Argon High 1519
- Arena instruction followingArena, 389 models
Top scoreGemini 4 Argon High 1529
Data analysis
Reasoning
Image generation
Each test’s results are its publisher’s own, credited on its page with its licence. A test page lists every model once, at its best configuration.