All makers

AI model maker

Google

79 models on Artificials, 41 with published test results and 17 with open weights you can download.

RSS feed
Best on the capability index157.3Gemini 3.7 Flash, #16 of 274
Top score on8 testsof the 28 its models are tested on
Newest releaseNano Banana 2.1Oct 6, 2026
Models7941 tested, 17 with open weights

Where Google stands

Its best model’s rank among all models tested on each skill area’s main test, the same tests model pages use. A longer bar means a better rank.

  • OverallGemini 3.7 Flash on Epoch Capabilities Index
    Top 6%
  • CodingGemini 3.7 Flash (high) on LiveBench coding
    Top 42%
  • MathGemini 3.8 Flash (high) on Mock AIME 2024–2025
    Top 9%
  • KnowledgeGemini 3.1 Pro (high) on SimpleQA Verified
    Top 4%
  • AgentsGemini 3.7 Flash (high) on LiveBench agentic coding
    Top 29%
  • LanguageGemini 3.8 Flash (high) on LiveBench language
    Top 10%
  • Data analysisGemini 3.1 Pro on LiveBench data analysis
    Top 34%
  • ReasoningGemini 3.8 Flash (high) on GPQA Diamond
    Top 2%
  • Image generationNano Banana 2 on T2I-CoReBench
    #1 of 40

On the capability index

Its best 12 of 35 models on the Epoch Capabilities Index, each with its rank among all 274 models. The top score is 167.3, held by Claude Opus 5.5 from Anthropic.

Every test

Its best result on each published test, and that result’s rank among every model tested. Scores from different tests are never combined.

TestBest modelScoreRank
Epoch Capabilities IndexGemini 3.7 Flash157.3#16 of 274
Arena hard promptsGemini 4 Argon High1551#1 of 389
Arena instruction followingGemini 4 Argon High1529#1 of 389
Arena textGemini 4 Argon High1525#1 of 389
Arena creative writingGemini 4 Argon High1519#1 of 387
Arena codingGemini 4 Argon High1560#1 of 384
Arena mathGemini 4 Argon High1530#1 of 372
GPQA DiamondGemini 3.8 Flash (high)95.4%#3 of 222
LiveBench instructionsGemini 3.8 Flash (high)81.4%#1 of 63
T2I-CoReBenchNano Banana 285.3%#1 of 40
Chess puzzlesGemini 3.8 Flash (high)61.0%#4 of 141
SimpleQA VerifiedGemini 3.1 Pro (high)73.5%#3 of 77
SWE-bench Verified, bash onlyGemini 3 Flash Preview (high)75.8%#2 of 42
Arena WebDevGemini 4 Argon High1680#8 of 123
DeepSWEGemini 3.8 Flash (high)73.8%#2 of 26
Mock AIME 2024–2025Gemini 3.8 Flash (high)98.9%#16 of 198
SWE-bench VerifiedGemini 3.5 Flash (high)79.3%#3 of 32
LiveBench languageGemini 3.8 Flash (high)87.8%#6 of 63
MATH Level 5Gemini 2.5 Pro Preview 050695.9%#10 of 97
FrontierMath Tier 4Gdm AI Co Mathematician75.6%#10 of 63
LiveBench overallGemini 3.7 Flash (high)78.8%#14 of 63
LiveBench reasoningGemini 3.8 Flash (high)89.3%#16 of 63
LiveBench mathematicsGemini 3.7 Flash (high)93.5%#17 of 63
FrontierMath Tiers 1–3Gemini 3.7 Flash (high)71.6%#22 of 81
LiveBench agentic codingGemini 3.7 Flash (high)58.3%#18 of 63
LiveBench data analysisGemini 3.1 Pro78.5%#21 of 63
LiveBench codingGemini 3.7 Flash (high)78.9%#26 of 63
CursorBenchGemini 3.8 Flash (high)39.6%#11 of 14

Every model

Newest first. The price is per 1M tokens, blended three input tokens to one output token: Google’s own listing, or else the middle price across hosts.

ModelReleasedContextPriceCapability index
Nano Banana 2.1Oct 6, 2026131.1K tokens$3.00Not on the index
Gemini 3.8 FlashSep 2, 20261M tokens$1.50156.7
Gemini 3.7 FlashAug 13, 20261M tokens$1.50157.3
Gemini Flash LatestAug 13, 20261M tokens$1.50Not on the index
Gemma 4 26B A4B MeroMeroJul 29, 2026262.1K tokens$0.185Not on the index
Gemma 4 31B MeroMero v2Jul 29, 2026262.1K tokens$0.188Not on the index
Gemini 3.6 FlashJul 21, 20261M tokens$1.50154.3
Gemini 3.5 Flash LiteJul 21, 20261M tokens$0.85145.1
Gemini Flash-Lite LatestJul 21, 20261M tokens$0.85Not on the index
Gemini Omni Flash PreviewJun 30, 2026131.1K tokens$5.50Not on the index
Nano Banana 2 LiteJun 30, 202665.5K tokens$7.69Not on the index
Gemini 3.5 Live Translate PreviewJun 9, 202616.4K tokens$7.88Not on the index
Gemma 4 12B ITMay 31, 2026262.1K tokens$0.15Not on the index
Nano Banana 2May 28, 2026131.1K tokens$15.38Not on the index
Nano Banana ProMay 28, 202665.5K tokens$31.50Not on the index

Models, prices, context windows and release dates from models.dev (MIT). Test results from LiveBench (Apache-2.0), Epoch AI (CC BY 4.0), T2I-CoReBench (CC BY-SA 4.0), Datacurve DeepSWE, compiled by Epoch AI (CC BY 4.0), Cursor CursorBench, compiled by Epoch AI (CC BY 4.0), Arena (CC BY 4.0) and SWE-bench (CC BY-NC 4.0). A model counts here when Google sells it, several hosts list it or a test covers it. Each figure is the source’s own; nothing is combined into a new score.