Compare models
HEAD TO HEAD

MiniMax-M3 vs Qwen3.8 Max

Published test results, prices, limits and features, side by side. Share the link to show someone exactly this comparison.

  1. MiniMax-M3MiniMax

At a glance

  • Capability index

    • MiniMax-M3146.9
    • Qwen3.8 Max156.4

    Qwen3.8 Max is 9.5 points higher.

  • Price per 1M tokens

    • MiniMax-M3$0.525
    • Qwen3.8 Max$3.00

    MiniMax-M3 is 5.7 times cheaper.

  • Context window

    • MiniMax-M31M tokens
    • Qwen3.8 Max1M tokens

    They take in about the same.

  • Released

    • MiniMax-M3Jun 1, 2026
    • Qwen3.8 MaxAug 3, 2026

    Qwen3.8 Max is 2 months newer.

Where each scores higher

On the 18 tests both have taken, each on its own scale. A gap of a point or two can sit within a test’s margin of error.

MiniMax-M3

No shared test where it scores higher.

Qwen3.8 Max

  • LiveBench overall+11.2 points
  • Arena text+42 points
  • LiveBench coding+4.7 points
  • Arena WebDev+189 points
  • Arena coding+29 points
  • LiveBench mathematics+14.4 points
  • Mock AIME 2024–2025+28.3 points
  • Arena math+66 points
  • LiveBench agentic coding+24.0 points
  • LiveBench language+2.8 points
  • LiveBench instructions+16.6 points
  • Arena creative writing+62 points
  • Arena instruction following+39 points
  • LiveBench data analysis+2.2 points
  • LiveBench reasoning+13.7 points
  • GPQA Diamond+1.8 points
  • Chess puzzles+15.0 points
  • Arena hard prompts+41 points

4 more tests have results for only one of them; the table below lists every result.

Everything side by side

Comparison of MiniMax-M3 vs Qwen3.8 Max
MeasureMiniMax-M3MiniMaxQwen3.8 MaxAlibaba
Model
MakerMiniMaxAlibaba
Sold here byMiniMax (minimax.io)Alibaba
API model IDMiniMax-M3qwen3.8-max
ReleasedJun 1, 2026Aug 3, 2026
Knowledge cutoffJan 2025Not reported
WeightsOpen: downloadableClosed: hosted access only
Developer’s countryChinaChina
Hosts selling it56including MiniMax directly36including Alibaba directly
Test results
Capability index146.9#68 of 274156.4 (best of these models)#22 of 274
LiveBench overall67.3%#61 of 6678.5% (best of these models)#17 of 66
Arena text1440#96 of 4131482 (best of these models)#23 of 413
LiveBench coding68.2%#64 of 6672.9% (best of these models)#52 of 66
Arena WebDev1482#58 of 1381671 (best of these models)#10 of 138
Arena coding1494#86 of 4081524 (best of these models)#28 of 408
LiveBench mathematics77.0%#65 of 6691.3% (best of these models)#27 of 66
Mock AIME 2024–202571.1%#133 of 29799.4% (best of these models)#14 of 297
Arena math1432#102 of 3961498 (best of these models)#18 of 396
FrontierMath Tiers 1–3Not tested74.7%#18 of 114
FrontierMath Tier 4Not tested46.3%#25 of 70
SimpleQA VerifiedNot tested45.8%#37 of 86
LiveBench agentic coding40.7%#60 of 6664.7% (best of these models)#7 of 66
DeepSWENot tested57.5%#36 of 69
LiveBench language76.8%#45 of 6679.7% (best of these models)#34 of 66
LiveBench instructions57.5%#61 of 6674.1% (best of these models)#13 of 66
Arena creative writing1406#99 of 4111468 (best of these models)#20 of 411
Arena instruction following1434#88 of 4131474 (best of these models)#31 of 413
LiveBench data analysis76.2%#34 of 6678.4% (best of these models)#24 of 66
LiveBench reasoning74.5%#61 of 6688.2% (best of these models)#23 of 66
GPQA Diamond90.9%#34 of 31992.7% (best of these models)#24 of 319
Chess puzzles14.0%#113 of 22729.0% (best of these models)#54 of 227
Arena hard prompts1463#90 of 4131503 (best of these models)#27 of 413
Price per million tokens
Input$0.30 (best of these models)$2.00
Output$1.20 (best of these models)$6.00
Blended, 3 input to 1 output$0.525 (best of these models)$3.00
Cached input$0.06 (best of these models)$0.25
Writing to the cacheNot reported$2.50
Long requests$0.60 in, $2.40 outabove 512K tokensNo separate price reported
The maker’s own priceThis listingThis listing
Limits
Context window1M tokens1M tokens
Max input512K tokens991K tokens (best of these models)
Max output512K tokens (best of these models)131.1K tokens
Features
Readsimages, text and videoimages, PDFs, text and video
Producestexttext
ReasoningYeseffort low, medium, high; can be switched off; thinking budget can be setYeseffort low, medium, xhigh; can be switched off; thinking budget can be set
Tool callingYesYes
Structured outputYes, such as JSONYes, such as JSON
File attachmentsYesYes
Temperature settingSupportedSupported

Bold marks the better value in each row. Prices are each model’s own list price, or the middle price across hosts when the maker does not sell it, blended as three input tokens for every output token. Test results as published by LiveBench (Apache-2.0), Epoch AI (CC BY 4.0), Datacurve DeepSWE, compiled by Epoch AI (CC BY 4.0), Arena (CC BY 4.0); model details and prices from models.dev (MIT). Confirm pricing with the provider before use.