AI model maker
Microsoft
8 models on Artificials, 1 with published test results and 8 with open weights you can download.
Where Microsoft stands
Its best model’s rank among all models tested on each skill area’s main test, the same tests model pages use. A longer bar means a better rank.
- OverallPhi-4 on Epoch Capabilities Index#157 of 274
- CodingPhi-4 on Arena coding#280 of 384
- MathPhi-4 on Mock AIME 2024–2025#145 of 198
- LanguagePhi-4 on Arena creative writing#292 of 387
- ReasoningPhi-4 on GPQA Diamond#132 of 222
On the capability index
Its 6 models on the Epoch Capabilities Index, each with its rank among all 274 models. The top score is 167.3, held by Claude Opus 5.5 from Anthropic.
Every test
Its best result on each published test, and that result’s rank among every model tested. Scores from different tests are never combined.
| Test | Best model | Score | Rank |
|---|---|---|---|
| Epoch Capabilities Index | Phi-4 | 130.4 | #157 of 274 |
| MATH Level 5 | Phi-4 | 64.9% | #39 of 97 |
| GPQA Diamond | Phi-4 | 56.1% | #132 of 222 |
| Arena math | Phi-4 | 1265 | #257 of 372 |
| Arena hard prompts | Phi-4 | 1278 | #282 of 389 |
| Arena coding | Phi-4 | 1306 | #280 of 384 |
| Mock AIME 2024–2025 | Phi-4 | 13.8% | #145 of 198 |
| Arena instruction following | Phi-4 | 1246 | #287 of 389 |
| Arena text | Phi-4 | 1256 | #293 of 389 |
| Arena creative writing | Phi-4 | 1210 | #292 of 387 |
| Chess puzzles | Phi-4 | 1.0% | #111 of 141 |
Every model
Newest first. The price is per 1M tokens, blended three input tokens to one output token: Microsoft’s own listing, or else the middle price across hosts.
| Model | Released | Context | Price | Capability index |
|---|---|---|---|---|
| Phi 4 Multimodal | Jul 26, 2025 | 128K tokens | $0.08 | Not on the index |
| Phi-4 | Dec 11, 2024 | 128K tokens | $0.219 | 130.4 |
| Phi-4-mini | Dec 11, 2024 | 128K tokens | $0.131 | Not on the index |
| Phi-4-mini-reasoning | Dec 11, 2024 | 128K tokens | $0.131 | Not on the index |
| Phi-4-multimodal | Dec 11, 2024 | 128K tokens | $0.14 | Not on the index |
| Phi-4-reasoning | Dec 11, 2024 | 32K tokens | $0.219 | Not on the index |
| Phi-4-reasoning-plus | Dec 11, 2024 | 32K tokens | $0.219 | Not on the index |
| Phi 4 Mini | Dec 1, 2024 | 131.1K tokens | $0.298 | Not on the index |
Models, prices, context windows and release dates from models.dev (MIT). Test results from Epoch AI (CC BY 4.0) and Arena (CC BY 4.0). A model counts here when Microsoft sells it, several hosts list it or a test covers it. Each figure is the source’s own; nothing is combined into a new score.