What changed, day by day.
New releases, price changes, new leaders on each test and models joining the capability index, newest first.
- New release
Mistral Large 4 was released by Mistral AI.
- New release
Mistral Large 4 (0) was released by Mistral AI.
- New release
Nano Banana 2.1 was released by Google.
- Joined the index
GPT-6 Luna joined the capability index at 156.3.
- Joined the index
GPT-6 Sol joined the capability index at 162.7.
- Joined the index
GPT-6.1 Sol joined the capability index at 166.1.
- Joined the index
Grok 4.7 joined the capability index at 153.5.
- Price change+82%
DeepSeek’s listed price for DeepSeek V4 Pro went from $0.435 in and $0.87 out to $0.66 in and $1.98 out per 1M tokens.
- New release
Grok Imagine Video 1.5 Lite was released by xAI.
- New leader
Gemini 4 Argon High took the top spot on Arena text at 1525, from Claude Opus 5.5.
- New leaderWithin margin
Gemini 4 Argon High took the top spot on Arena coding at 1559, from Claude Opus 4.6.
- New leaderWithin margin
Gemini 4 Argon High took the top spot on Arena math at 1528, from Claude Fable 5.
- New leader
Gemini 4 Argon High took the top spot on Arena hard prompts at 1550, from Claude Opus 5.5.
- New leaderWithin margin
Gemini 4 Argon High took the top spot on Arena creative writing at 1522, from Claude Opus 5.5.
- New leader
Gemini 4 Argon High took the top spot on Arena instruction following at 1529, from Claude Opus 5.5.
- New leader
GPT-6.1 Sol took the top spot on FrontierMath Tier 4 at 100.0%, from GPT-6 Astra.
- Joined the index
Claude Opus 5.5 joined the capability index at 167.3.
- Joined the index
Claude Sonnet 5.5 joined the capability index at 165.2.
- New release
GPT-6.1 Sol was released by OpenAI.
- New release
GPT-6.1 Sol Pro was released by OpenAI.
- New leader
Claude Opus 5.5 took the top spot on CursorBench at 57.8%, from Claude Fable 5.1.
- New release
Claude Sonnet 5.5 was released by Anthropic.
- New leader
Claude Sonnet 5.5 took the top spot on LiveBench coding at 91.4%, from Claude Opus 5.5.
- Price change−38%
Alibaba’s listed price for Qwen3.7 Plus went from $0.50 in and $3.00 out to $0.40 in and $1.60 out per 1M tokens.
- New release
MiniMax-M3.1-Flash-Preview was released by MiniMax.
- Joined the index
amazon--nova-pro joined the capability index at 123.8.
- Joined the index
Llama 3 70B Instruct joined the capability index at 122.9.
- Joined the index
Llama 3 8B Instruct joined the capability index at 116.3.
- New leaderWithin margin
Claude Opus 4.6 took the top spot on Arena coding at 1551, from Claude Fable 5.
- New leaderWithin margin
Claude Opus 5.5 took the top spot on Arena text at 1509, from Claude Fable 5.
- New leaderWithin margin
Claude Opus 5.5 took the top spot on Arena hard prompts at 1541, from Claude Opus 4.6.
- New leaderWithin margin
Claude Opus 5.5 took the top spot on Arena creative writing at 1521, from Claude Fable 5.
- New leaderWithin margin
Claude Opus 5.5 took the top spot on Arena instruction following at 1516, from Claude Opus 4.6.
- New release
LongCat 2.5 Preview was released by Meituan.
- New release
LongCat 2.5 Preview Free was released by Meituan.
- New release
MiMo V2.6 Flash Uncensored was released by Xiaomi.
- New release
GLM 5.3 Prime was released by Z.ai.
- New release
Qwen 3.8 Max Prime was released by Alibaba.
- New release
Solar Mini 4 was released by Upstage.
- New release
Claude Opus 5.5 was released by Anthropic.
- New release
Claude Opus 5.5 Fast was released by Anthropic.
- New release
Command A+ was released by Cohere.
- New release
GPT-6 Luna was released by OpenAI.
- New release
GPT-6 Luna Pro was released by OpenAI.
- New release
GPT-6 Sol was released by OpenAI.
- New release
GPT-6 Sol Pro was released by OpenAI.
- New release
MiMo-V2.6-Flash was released by Xiaomi.
- New release
MiMo-V2.6-Pro was released by Xiaomi.
- New leader
Claude Opus 5.5 took the top spot on LiveBench coding at 89.3%, from Claude Fable 5.1.
- New leader
Claude Opus 5.5 took the top spot on LiveBench mathematics at 97.1%, from Claude Fable 5.1.
- New release
Grok 4.7 was released by xAI.
- New release
MiMo-V2.6-Pro-UltraSpeed was released by Xiaomi.
- Price change+100%
Z.ai’s listed price for GLM-5.3-Flash went from $0.075 in and $0.25 out to $0.15 in and $0.50 out per 1M tokens.
- New release
GLM-5.3-FlashX was released by Z.ai.
- New release
Qwen3.8 Omni Flash was released by Alibaba.
- New release
Step 5 Preview was released by StepFun.
- New release
GPT Astra Latest was released by OpenAI.
- New release
GPT Luna Latest was released by OpenAI.
- New release
GPT Terra Latest was released by OpenAI.
- New release
kimi-for-coding was released by Moonshot AI.
- New release
DeepSeek Flash Latest was released by DeepSeek.
- New release
DeepSeek V4 Flash was released by DeepSeek.
- New release
DeepSeek V4 Flash Vision Exp was released by DeepSeek.
- New release
DeepSeek V4.1 Flash was released by DeepSeek.
- New release
deepseek-ai/DeepSeek-V4.1-Flash-Fast was released by DeepSeek.
- New release
North Small Translate was released by Cohere.
- New release
GPT Image 2.5 Flare was released by OpenAI.
- New release
GPT Image 2.5 Sunburst was released by OpenAI.
- New release
Mercury 2.5 was released by Inception.
Releases and listed prices from models.dev (MIT); LiveBench (Apache-2.0); Epoch AI (CC BY 4.0); Cursor CursorBench, compiled by Epoch AI (CC BY 4.0); Arena (CC BY 4.0). A price change is the maker’s own listing changing, not the middle price across hosts; a percentage compares prices blended as three input tokens for every output token. “Within margin” marks a new leader whose lead is inside the range its publisher gives for its score: the next model could still pass it as more votes come in. Changes count from when the site began keeping this record, so older ones are not listed. Each week on one page. RSS feed