AI model maker
Meta
34 models on Artificials, 11 with published test results and 34 with open weights you can download.
Where Meta stands
Its best model’s rank among all models tested on each skill area’s main test, the same tests model pages use. A longer bar means a better rank.
- OverallMuse Spark 1.3 on Epoch Capabilities IndexTop 7%
- CodingMuse Spark 1.3 (max) on Arena codingTop 3%
- MathMuse Spark 1.3 (xhigh) on Mock AIME 2024–2025Top 8%
- KnowledgeMuse Spark 1.2 (xhigh) on SimpleQA VerifiedTop 20%
- LanguageMuse Spark on Arena creative writingTop 5%
- ReasoningMuse Spark on GPQA DiamondTop 20%
On the capability index
Its best 12 of 22 models on the Epoch Capabilities Index, each with its rank among all 274 models. The top score is 167.3, held by Claude Opus 5.5 from Anthropic.
Every test
Its best result on each published test, and that result’s rank among every model tested. Scores from different tests are never combined.
| Test | Best model | Score | Rank |
|---|---|---|---|
| Epoch Capabilities Index | Muse Spark 1.3 | 156.8 | #19 of 274 |
| Arena hard prompts | Muse Spark 1.3 (max) | 1517 | #8 of 389 |
| Arena text | Muse Spark 1.3 (max) | 1494 | #8 of 389 |
| Arena math | Muse Spark 1.3 (max) | 1509 | #8 of 372 |
| Arena coding | Muse Spark 1.3 (max) | 1539 | #9 of 384 |
| Arena instruction following | Muse Spark 1.3 (max) | 1486 | #13 of 389 |
| Arena creative writing | Muse Spark | 1466 | #18 of 387 |
| Mock AIME 2024–2025 | Muse Spark 1.3 (xhigh) | 99.2% | #14 of 198 |
| Arena WebDev | Muse Spark 1.3 (max) | 1657 | #12 of 123 |
| Chess puzzles | Muse Spark 1.3 (max) | 38.0% | #25 of 141 |
| SimpleQA Verified | Muse Spark 1.2 (xhigh) | 60.3% | #15 of 77 |
| GPQA Diamond | Muse Spark | 89.8% | #44 of 222 |
| FrontierMath Tiers 1–3 | Muse Spark 1.3 (xhigh) | 74.4% | #19 of 81 |
| FrontierMath Tier 4 | Muse Spark 1.3 (max) | 46.3% | #19 of 63 |
| MATH Level 5 | Llama 4 Maverick 17B 128E Instruct FP8 | 73.0% | #33 of 97 |
| CursorBench | Muse Spark 1.3 (max) | 41.6% | #8 of 14 |
| DeepSWE | Muse Spark 1.2 (xhigh) | 54.9% | #16 of 26 |
| SWE-bench Verified, bash only | Llama 4 Maverick Instruct | 21.0% | #39 of 42 |
Every model
Newest first. The price is per 1M tokens, blended three input tokens to one output token: Meta’s own listing, or else the middle price across hosts.
| Model | Released | Context | Price | Capability index |
|---|---|---|---|---|
| Llama 4 Scout 17B 16E Instruct | Apr 5, 2025 | 128K tokens | $0.345 | 129.6 |
| Llama 4 Maverick 17B Instruct | Apr 5, 2025 | 1M tokens | $0.415 | Not on the index |
| Llama 4 Maverick 17B Instruct | Apr 5, 2025 | 1M tokens | $0.423 | Not on the index |
| Llama 4 Scout 17B Instruct | Apr 5, 2025 | 131.1K tokens | $0.283 | Not on the index |
| Llama 4 Scout 17B Instruct | Apr 5, 2025 | 10M tokens | $0.293 | Not on the index |
| Llama Guard 4 12B | Apr 5, 2025 | 163.8K tokens | $0.18 | Not on the index |
| Llama 4 Maverick 17b 128e Instruct | Apr 1, 2025 | 128K tokens | No paid listing | 132.2 |
| Llama 4 Maverick 17B 128E Instruct FP8 | Jan 15, 2025 | 1M tokens | $0.415 | Not on the index |
| Llama 4 Maverick | Jan 1, 2025 | 128K tokens | $0.304 | Not on the index |
| Llama 4 Scout | Jan 1, 2025 | 1.3M tokens | $0.15 | Not on the index |
| Llama 3.3 70B | Dec 6, 2024 | 128K tokens | $1.23 | Not on the index |
| Llama 3.3 70B | Dec 6, 2024 | 131.1K tokens | $2.00 | Not on the index |
| Llama 3.3 70B Instruct | Dec 6, 2024 | 128K tokens | $0.72 | Not on the index |
| Llama 3.3 70B Turbo | Dec 6, 2024 | 131.1K tokens | $0.155 | Not on the index |
| Llama 3.3 70B Versatile | Dec 6, 2024 | 131.1K tokens | $0.64 | Not on the index |
Models, prices, context windows and release dates from models.dev (MIT). Test results from Epoch AI (CC BY 4.0), Datacurve DeepSWE, compiled by Epoch AI (CC BY 4.0), Cursor CursorBench, compiled by Epoch AI (CC BY 4.0), Arena (CC BY 4.0) and SWE-bench (CC BY-NC 4.0). A model counts here when Meta sells it, several hosts list it or a test covers it. Each figure is the source’s own; nothing is combined into a new score.