All makers

AI model maker

xAI

33 models on Artificials, 16 with published test results.

RSS feed
Best on the capability index156.4Grok 4.6, #21 of 274
Best rankTop 5%Grok 4.20 Beta1 on Arena creative writing
Newest releaseGrok Imagine Video 1.5 LiteOct 1, 2026
Models3316 tested, 0 with open weights

Where xAI stands

Its best model’s rank among all models tested on each skill area’s main test, the same tests model pages use. A longer bar means a better rank.

  • OverallGrok 4.6 on Epoch Capabilities Index
    Top 8%
  • CodingGrok 4.7 (xhigh) on LiveBench coding
    #35 of 63
  • MathGrok 4.6 (xhigh) on Mock AIME 2024–2025
    Top 8%
  • KnowledgeGrok 4.7 (xhigh) on SimpleQA Verified
    Top 24%
  • AgentsGrok 4.6 on LiveBench agentic coding
    Top 34%
  • LanguageGrok 4.6 on LiveBench language
    Top 27%
  • Data analysisGrok 4.7 (xhigh) on LiveBench data analysis
    Top 45%
  • ReasoningGrok 4.6 (high) on GPQA Diamond
    Top 5%

On the capability index

Its 10 models on the Epoch Capabilities Index, each with its rank among all 274 models. The top score is 167.3, held by Claude Opus 5.5 from Anthropic.

Every test

Its best result on each published test, and that result’s rank among every model tested. Scores from different tests are never combined.

TestBest modelScoreRank
Epoch Capabilities IndexGrok 4.6156.4#21 of 274
Arena creative writingGrok 4.20 Beta11463#19 of 387
GPQA DiamondGrok 4.6 (high)94.0%#11 of 222
Mock AIME 2024–2025Grok 4.6 (xhigh)99.2%#14 of 198
Arena textGrok 4.20 Beta11475#31 of 389
Arena instruction followingGrok 4.51461#40 of 389
Arena WebDevGrok 4.7 (xhigh)1638#13 of 123
Arena codingGrok 4.51514#41 of 384
Arena mathGrok 4.51473#40 of 372
Arena hard promptsGrok 4.51489#42 of 389
Chess puzzlesGrok 4.6 (high)40.0%#20 of 141
LiveBench instructionsGrok 4.7 (xhigh)75.3%#10 of 63
LiveBench reasoningGrok 4.690.5%#10 of 63
MATH Level 5Grok 3 Mini Beta (low)90.9%#16 of 97
LiveBench mathematicsGrok 4.7 (xhigh)95.7%#12 of 63
SimpleQA VerifiedGrok 4.7 (xhigh)56.0%#18 of 77
LiveBench overallGrok 4.678.0%#16 of 63
LiveBench languageGrok 4.683.7%#17 of 63
FrontierMath Tiers 1–3Grok 4.6 (xhigh)66.0%#27 of 81
LiveBench agentic codingGrok 4.657.0%#21 of 63
DeepSWEGrok 4.6 (medium)67.5%#9 of 26
CursorBenchGrok 4.7 (xhigh)46.3%#5 of 14
FrontierMath Tier 4Grok 4.6 (xhigh)31.7%#26 of 63
LiveBench data analysisGrok 4.7 (xhigh)76.9%#28 of 63
LiveBench codingGrok 4.7 (xhigh)77.2%#35 of 63

Every model

Newest first. The price is per 1M tokens, blended three input tokens to one output token: xAI’s own listing, or else the middle price across hosts.

ModelReleasedContextPriceCapability index
Grok Imagine Video 1.5 LiteOct 1, 20261K tokensNo paid listingNot on the index
Grok 4.7Sep 21, 2026500K tokens$3.00153.5
Grok 4.6Aug 12, 2026500K tokens$3.00156.4
Grok Imagine Image 2.0Aug 7, 202666K tokensNo paid listingNot on the index
Grok 4.5Jul 8, 2026500K tokens$3.00153.9
Grok Imagine Video 1.5May 30, 20261K tokensNo paid listingNot on the index
Grok LatestMay 3, 2026500K tokens$3.00Not on the index
Grok 4.3Apr 17, 20261M tokens$1.56149.2
Grok Build 0.1Apr 16, 2026256K tokens$1.25Not on the index
Grok Imagine Image QualityApr 3, 202616K tokensNo paid listingNot on the index
Grok 4.20 Multi-AgentMar 10, 20262M tokens$1.56Not on the index
Grok 4.20Mar 9, 20262M tokens$1.77152.0
Grok 4.20 (0309 non reasoning)Mar 9, 20261M tokens$1.56Not on the index
Grok 4.20 (0309 reasoning)Mar 9, 20261M tokens$1.56Not on the index
Grok 4.20 (beta 0309 non reasoning)Mar 9, 20262M tokens$3.00Not on the index

Models, prices, context windows and release dates from models.dev (MIT). Test results from LiveBench (Apache-2.0), Epoch AI (CC BY 4.0), Datacurve DeepSWE, compiled by Epoch AI (CC BY 4.0), Cursor CursorBench, compiled by Epoch AI (CC BY 4.0) and Arena (CC BY 4.0). A model counts here when xAI sells it, several hosts list it or a test covers it. Each figure is the source’s own; nothing is combined into a new score.