Compare models
HEAD TO HEAD

Grok 4.6 vs Muse Spark 1.3

Published test results, prices, limits and features, side by side. Share the link to show someone exactly this comparison.

At a glance

  • Capability index

    • Grok 4.6156.4
    • Muse Spark 1.3156.8

    Muse Spark 1.3 is 0.3 points higher.

  • Price per 1M tokens

    • Grok 4.6$3.00
    • Muse Spark 1.3$2.00

    Muse Spark 1.3 is 33% cheaper.

  • Context window

    • Grok 4.6500K tokens
    • Muse Spark 1.31M tokens

    Muse Spark 1.3 takes in 2.1 times as much.

  • Released

    • Grok 4.6Aug 12, 2026
    • Muse Spark 1.3Sep 2, 2026

    Muse Spark 1.3 is 21 days newer.

Where each scores higher

On the 20 tests both have taken, each on its own scale. A gap of a point or two can sit within a test’s margin of error.

Grok 4.6

  • LiveBench language+0.9 points
  • LiveBench reasoning+0.9 points
  • Chess puzzles+2.0 points

Muse Spark 1.3

  • LiveBench overall+3.5 points
  • Arena text+41 points
  • LiveBench coding+4.3 points
  • Arena WebDev+37 points
  • Arena coding+31 points
  • LiveBench mathematics+3.4 points
  • FrontierMath Tiers 1–3+8.4 points
  • FrontierMath Tier 4+14.6 points
  • Arena math+61 points
  • LiveBench agentic coding+7.1 points
  • CursorBench+0.2 points
  • LiveBench instructions+6.1 points
  • Arena creative writing+14 points
  • Arena instruction following+35 points
  • LiveBench data analysis+5.7 points
  • Arena hard prompts+38 points

Level on Mock AIME 2024–2025.

3 more tests have results for only one of them; the table below lists every result.

Everything side by side

Comparison of Grok 4.6 vs Muse Spark 1.3
MeasureGrok 4.6xAIMuse Spark 1.3Meta AI
Model
MakerxAIMeta AI
Sold here byxAIDevPass (LLM Gateway)
API model IDgrok-4.6muse-spark-1.3
ReleasedAug 12, 2026Sep 2, 2026
Knowledge cutoffFeb 1, 2026Not reported
WeightsClosed: hosted access onlyClosed: hosted access only
Developer’s countryUnited StatesUnited States
Hosts selling it35including xAI directly11
Test results
Capability index156.4#21 of 274156.8 (best of these models)#19 of 274
LiveBench overall78.0%#18 of 6681.6% (best of these models)#7 of 66
Arena text1454#73 of 4131494 (best of these models)#9 of 413
LiveBench coding76.8%#42 of 6681.1% (best of these models)#17 of 66
Arena WebDev1620#23 of 1381657 (best of these models)#14 of 138
Arena coding1507#64 of 4081539 (best of these models)#11 of 408
LiveBench mathematics92.6%#25 of 6696.0% (best of these models)#12 of 66
Mock AIME 2024–202599.2%#15 of 29799.2%#15 of 297
FrontierMath Tiers 1–366.0%#28 of 11474.4% (best of these models)#19 of 114
FrontierMath Tier 431.7%#33 of 7046.3% (best of these models)#25 of 70
Arena math1448#78 of 3961509 (best of these models)#9 of 396
SimpleQA Verified49.3%#27 of 86Not tested
LiveBench agentic coding57.0%#22 of 6664.1% (best of these models)#9 of 66
DeepSWE67.5%#20 of 69Not tested
CursorBench41.4%#24 of 6241.6% (best of these models)#22 of 62
LiveBench language83.7% (best of these models)#19 of 6682.8%#24 of 66
LiveBench instructions71.9%#20 of 6678.0% (best of these models)#4 of 66
Arena creative writing1445#52 of 4111459 (best of these models)#28 of 411
Arena instruction following1451#63 of 4131486 (best of these models)#16 of 413
LiveBench data analysis73.9%#40 of 6679.6% (best of these models)#12 of 66
LiveBench reasoning90.5% (best of these models)#12 of 6689.7%#15 of 66
GPQA Diamond94.0%#11 of 319Not tested
Chess puzzles40.0% (best of these models)#24 of 22738.0%#32 of 227
Arena hard prompts1479#67 of 4131517 (best of these models)#10 of 413
Price per million tokens
Input$2.00$1.25 (best of these models)
Output$6.00$4.25 (best of these models)
Blended, 3 input to 1 output$3.00$2.00 (best of these models)
Cached input$0.50$0.15 (best of these models)
Long requests$4.00 in, $12.00 outabove 200K tokensNo separate price reported
The maker’s own priceThis listingNot sold directly
Limits
Context window500K tokens1M tokens (best of these models)
Max input372K tokens1M tokens (best of these models)
Max output500K tokens (best of these models)131.1K tokens
Features
Readsimages, PDFs and textaudio, images, PDFs, text and video
Producestexttext
ReasoningYeseffort low, medium, high, xhighYeseffort minimal, low, medium, high, xhigh
Tool callingYesYes
Structured outputYes, such as JSONYes, such as JSON
File attachmentsYesYes
Temperature settingSupportedSupported

Bold marks the better value in each row. Prices are each model’s own list price, or the middle price across hosts when the maker does not sell it, blended as three input tokens for every output token. Test results as published by LiveBench (Apache-2.0), Epoch AI (CC BY 4.0), Datacurve DeepSWE, compiled by Epoch AI (CC BY 4.0), Cursor CursorBench, compiled by Epoch AI (CC BY 4.0), Arena (CC BY 4.0); model details and prices from models.dev (MIT). Confirm pricing with the provider before use.