All makers

AI model maker

Meta

34 models on Artificials, 11 with published test results and 34 with open weights you can download.

RSS feed
Best on the capability index156.8Muse Spark 1.3, #19 of 274
Best rankTop 3%Muse Spark 1.3 (max) on Arena hard prompts
Newest releaseLlama 4 Scout 17B 16E InstructApr 5, 2025
Models3411 tested, 34 with open weights

Where Meta stands

Its best model’s rank among all models tested on each skill area’s main test, the same tests model pages use. A longer bar means a better rank.

  • OverallMuse Spark 1.3 on Epoch Capabilities Index
    Top 7%
  • CodingMuse Spark 1.3 (max) on Arena coding
    Top 3%
  • MathMuse Spark 1.3 (xhigh) on Mock AIME 2024–2025
    Top 8%
  • KnowledgeMuse Spark 1.2 (xhigh) on SimpleQA Verified
    Top 20%
  • LanguageMuse Spark on Arena creative writing
    Top 5%
  • ReasoningMuse Spark on GPQA Diamond
    Top 20%

On the capability index

Its best 12 of 22 models on the Epoch Capabilities Index, each with its rank among all 274 models. The top score is 167.3, held by Claude Opus 5.5 from Anthropic.

Every test

Its best result on each published test, and that result’s rank among every model tested. Scores from different tests are never combined.

TestBest modelScoreRank
Epoch Capabilities IndexMuse Spark 1.3156.8#19 of 274
Arena hard promptsMuse Spark 1.3 (max)1517#8 of 389
Arena textMuse Spark 1.3 (max)1494#8 of 389
Arena mathMuse Spark 1.3 (max)1509#8 of 372
Arena codingMuse Spark 1.3 (max)1539#9 of 384
Arena instruction followingMuse Spark 1.3 (max)1486#13 of 389
Arena creative writingMuse Spark1466#18 of 387
Mock AIME 2024–2025Muse Spark 1.3 (xhigh)99.2%#14 of 198
Arena WebDevMuse Spark 1.3 (max)1657#12 of 123
Chess puzzlesMuse Spark 1.3 (max)38.0%#25 of 141
SimpleQA VerifiedMuse Spark 1.2 (xhigh)60.3%#15 of 77
GPQA DiamondMuse Spark89.8%#44 of 222
FrontierMath Tiers 1–3Muse Spark 1.3 (xhigh)74.4%#19 of 81
FrontierMath Tier 4Muse Spark 1.3 (max)46.3%#19 of 63
MATH Level 5Llama 4 Maverick 17B 128E Instruct FP873.0%#33 of 97
CursorBenchMuse Spark 1.3 (max)41.6%#8 of 14
DeepSWEMuse Spark 1.2 (xhigh)54.9%#16 of 26
SWE-bench Verified, bash onlyLlama 4 Maverick Instruct21.0%#39 of 42

Every model

Newest first. The price is per 1M tokens, blended three input tokens to one output token: Meta’s own listing, or else the middle price across hosts.

ModelReleasedContextPriceCapability index
Llama 4 Scout 17B 16E InstructApr 5, 2025128K tokens$0.345129.6
Llama 4 Maverick 17B InstructApr 5, 20251M tokens$0.415Not on the index
Llama 4 Maverick 17B InstructApr 5, 20251M tokens$0.423Not on the index
Llama 4 Scout 17B InstructApr 5, 2025131.1K tokens$0.283Not on the index
Llama 4 Scout 17B InstructApr 5, 202510M tokens$0.293Not on the index
Llama Guard 4 12BApr 5, 2025163.8K tokens$0.18Not on the index
Llama 4 Maverick 17b 128e InstructApr 1, 2025128K tokensNo paid listing132.2
Llama 4 Maverick 17B 128E Instruct FP8Jan 15, 20251M tokens$0.415Not on the index
Llama 4 MaverickJan 1, 2025128K tokens$0.304Not on the index
Llama 4 ScoutJan 1, 20251.3M tokens$0.15Not on the index
Llama 3.3 70BDec 6, 2024128K tokens$1.23Not on the index
Llama 3.3 70BDec 6, 2024131.1K tokens$2.00Not on the index
Llama 3.3 70B InstructDec 6, 2024128K tokens$0.72Not on the index
Llama 3.3 70B TurboDec 6, 2024131.1K tokens$0.155Not on the index
Llama 3.3 70B VersatileDec 6, 2024131.1K tokens$0.64Not on the index

Models, prices, context windows and release dates from models.dev (MIT). Test results from Epoch AI (CC BY 4.0), Datacurve DeepSWE, compiled by Epoch AI (CC BY 4.0), Cursor CursorBench, compiled by Epoch AI (CC BY 4.0), Arena (CC BY 4.0) and SWE-bench (CC BY-NC 4.0). A model counts here when Meta sells it, several hosts list it or a test covers it. Each figure is the source’s own; nothing is combined into a new score.