All makers

AI model maker

Alibaba

130 models on Artificials, 51 with published test results and 78 with open weights you can download.

RSS feed
Best on the capability index156.4Qwen 3.8 Max, #22 of 274
Top score on1 testof the 27 its models are tested on
Newest releaseQwen 3.8 Max PrimeSep 23, 2026
Models13051 tested, 78 with open weights

Where Alibaba stands

Its best model’s rank among all models tested on each skill area’s main test, the same tests model pages use. A longer bar means a better rank.

  • OverallQwen 3.8 Max on Epoch Capabilities Index
    Top 9%
  • CodingQwen3.6 Plus on LiveBench coding
    Top 47%
  • MathQwen3.8 Max 0902 (xhigh) on Mock AIME 2024–2025
    #1 of 198
  • KnowledgeQwen3.7 Max on SimpleQA Verified
    Top 25%
  • AgentsQwen3.8 Max on LiveBench agentic coding
    Top 10%
  • LanguageQwen3.7 Max on LiveBench language
    Top 50%
  • Data analysisQwen3.8 Max on LiveBench data analysis
    Top 35%
  • ReasoningQwen3.8 Max (xhigh) on GPQA Diamond
    Top 10%
  • Image generationQwen Image 2512 on T2I-CoReBench
    Top 28%

On the capability index

Its best 12 of 41 models on the Epoch Capabilities Index, each with its rank among all 274 models. The top score is 167.3, held by Claude Opus 5.5 from Anthropic.

Every test

Its best result on each published test, and that result’s rank among every model tested. Scores from different tests are never combined.

TestBest modelScoreRank
Epoch Capabilities IndexQwen 3.8 Max156.4#22 of 274
Mock AIME 2024–2025Qwen3.8 Max 0902 (xhigh)100.0%#1 of 198
Arena mathQwen3.8 Max1498#16 of 372
Arena creative writingQwen3.8 Max1468#17 of 387
Arena textQwen3.8 Max1482#20 of 389
Arena hard promptsQwen3.8 Max1503#23 of 389
Arena codingQwen3.7 Max Preview1524#23 of 384
MATH Level 5qwen3-max-2025-09-2397.1%#6 of 97
Arena instruction followingQwen3.8 Max1474#25 of 389
Arena WebDevQwen3.8 Max1671#9 of 123
LiveBench instructionsQwen3.8 Flash Next77.1%#5 of 63
LiveBench agentic codingQwen3.8 Max64.7%#6 of 63
GPQA DiamondQwen3.8 Max (xhigh)92.7%#22 of 222
Chess puzzlesQwen3.8 Max 0902 (xhigh)40.0%#20 of 141
SWE-bench VerifiedQwen3.7 Max77.3%#7 of 32
FrontierMath Tiers 1–3Qwen3.8 Max (xhigh)74.7%#18 of 81
LiveBench overallQwen3.8 Max78.5%#15 of 63
SimpleQA VerifiedQwen3.7 Max55.8%#19 of 77
T2I-CoReBenchQwen Image 251262.4%#11 of 40
FrontierMath Tier 4Qwen3.8 Max (xhigh)46.3%#19 of 63
LiveBench reasoningQwen3.8 Max88.2%#21 of 63
LiveBench data analysisQwen3.8 Max78.4%#22 of 63
LiveBench mathematicsQwen3.8 Max91.3%#24 of 63
LiveBench codingQwen3.6 Plus78.2%#29 of 63
LiveBench languageQwen3.7 Max79.7%#31 of 63
DeepSWEQwen3.8 Max (xhigh)57.5%#15 of 26
SWE-bench Verified, bash onlyQwen3-Coder 480B-A35B Instruct55.4%#25 of 42

Every model

Newest first. The price is per 1M tokens, blended three input tokens to one output token: Alibaba’s own listing, or else the middle price across hosts.

ModelReleasedContextPriceCapability index
Qwen 3.8 Max PrimeSep 23, 20261M tokens$6.00Not on the index
Qwen3.8 Omni FlashSep 17, 20261M tokens$0.23Not on the index
Qwen3.8 Flash NextAug 27, 2026262.1K tokens$0.275Not on the index
Qwen3.8 FlashAug 26, 20261M tokens$0.23Not on the index
Qwen3.8 27BAug 14, 2026262.1K tokens$0.956149.4
Qwen3.8 27BAug 14, 2026131.1K tokens$1.12Not on the index
Qwen3.8 2.4T A95BAug 12, 2026262.1K tokens$3.00Not on the index
Qwen3.8 MaxAug 3, 20261M tokens$3.00156.4
Qwen3.8 Max 0902Aug 3, 20261M tokens$3.00155.1
Qwen 3.8 27B ObliteratedJul 29, 2026524.3K tokens$0.563Not on the index
Qwen 3.8 27B UncensoredJul 29, 2026524.3K tokens$0.413Not on the index
Qwen3.8 Max PreviewJul 19, 20261M tokens$0.507Not on the index
Qwen3.7 FlashJul 15, 20261M tokens$0.055144.6
Qwen3.7 PlusJun 2, 20261M tokens$0.70147.4
Wan2.7 ImageMay 29, 20268.2K tokensNo paid listingNot on the index

Models, prices, context windows and release dates from models.dev (MIT). Test results from LiveBench (Apache-2.0), Epoch AI (CC BY 4.0), T2I-CoReBench (CC BY-SA 4.0), Datacurve DeepSWE, compiled by Epoch AI (CC BY 4.0), Arena (CC BY 4.0) and SWE-bench (CC BY-NC 4.0). A model counts here when Alibaba sells it, several hosts list it or a test covers it. Each figure is the source’s own; nothing is combined into a new score.