MODEL COMPARISON

Compare AI models.

Capability, price and context for every AI model we track. Open any model for its scores, prices and providers.

At a glance

The leaders on each measure right now. Open any model for its full profile.

Capability

Claude Opus 5.5 (167.3) and GPT-6 Astra (166.4) are the most capable models, followed by GPT-6.1 Sol (166.1) and Claude Sonnet 5.5 (165.0).

Price

Llama 3-8B ($0.04) and Qwen3.7 Flash ($0.055) are the cheapest models with a capability score, followed by Llama 3.1-8B ($0.06) and Gemma 3 4B ($0.0625).

Context window

Gemini 2.0 Pro (2.1M) and Grok 4.20 (2M) take in the most text at once, followed by Grok 4 Fast (2M) and GPT-6 Astra (1.1M).

Open weights

Kimi K3 (157.4) and DeepSeek V4 Pro 0813 (155.3) are the most capable open-weights models, followed by DeepSeek V4.1 Flash (154.9) and DeepSeek V4 Flash 0731 (154.5).

New

Mistral Large 4 (Oct 6, 2026) and Nano Banana 2.1 (Oct 6, 2026) are the newest releases, followed by Mistral Large 4 (0) (Oct 6, 2026) and Grok Imagine Video 1.5 Lite (Oct 1, 2026).

Highlights

The ten most capable models, compared on score, price and context window.

Capability index

The ten highest scores. Higher is better.
  1. Claude Opus 5.5167.3
  2. GPT-6 Astra166.4
  3. GPT-6.1 Sol166.1
  4. Claude Sonnet 5.5165.0
  5. Claude Fable 5.1164.7
  6. Claude Opus 5162.8
  7. GPT-6 Sol162.7
  8. GPT-5.5 Pro162.1
  9. Claude Fable 5162.1
  10. GPT-5.6 Sol161.7

Blended price

Same ten models, USD per 1M tokens. Lower is cheaper.
  1. GPT-6.1 Sol$4.00
  2. Claude Sonnet 5.5$4.00
  3. GPT-6 Sol$4.00
  4. Claude Opus 5.5$8.00
  5. GPT-5.6 Sol$8.00
  6. Claude Opus 5$10.00
  7. GPT-6 Astra$20.00
  8. Claude Fable 5.1$20.00
  9. Claude Fable 5$20.00
  10. GPT-5.5 Pro$67.50

Context window

Same ten models, tokens. Higher takes in more.
  1. GPT-6 Astra1.1M
  2. GPT-6.1 Sol1.1M
  3. GPT-6 Sol1.1M
  4. GPT-5.5 Pro1.1M
  5. GPT-5.6 Sol1.1M
  6. Claude Opus 5.51M
  7. Claude Sonnet 5.51M
  8. Claude Fable 5.11M
  9. Claude Opus 51M
  10. Claude Fable 51M

Not on the index yet: MiMo-V2.6-Flash (Sep 22), MiMo-V2.6-Pro (Sep 22), Step 5 Preview (Sep 16), DeepSeek V4 Flash Vision Exp (Sep 10), Qwen3.8 Flash Next (Aug 27), Granite 4.2 8B (Aug 24). Epoch AI adds a model once enough of its tests have run; the results already published are on each model’s page.

New models

The newest releases from known makers. A capability score follows once a test covers the model.

  1. Mistral Large 4Mistral AIReleasedOct 6, 2026Capability indexNot tested yetBlended price$1.03Context window524.3K
  2. Nano Banana 2.1GoogleReleasedOct 6, 2026Capability indexNot tested yetBlended price$3.00Context window131.1K
  3. Mistral Large 4 (0)Mistral AIReleasedOct 6, 2026Capability indexNot tested yetBlended price$1.03Context window524.3K
  4. Grok Imagine Video 1.5 LitexAIReleasedOct 1, 2026Capability indexNot tested yetBlended priceNot listedContext window1K
  5. GPT-6.1 SolOpenAIReleasedSep 29, 2026Capability index166.1Blended price$4.00Context window1.1M
  6. GPT-6.1 Sol ProOpenAIReleasedSep 29, 2026Capability indexNot tested yetBlended price$4.00Context window1.1M
  7. Claude Sonnet 5.5AnthropicReleasedSep 28, 2026Capability index165.0Blended price$4.00Context window1M
  8. MiniMax-M3.1-Flash-PreviewMiniMaxReleasedSep 27, 2026Capability indexNot tested yetBlended priceNot listedContext window1M
  9. LongCat 2.5 PreviewMeituanReleasedSep 25, 2026Capability indexNot tested yetBlended price$0.525Context window1M
  10. LongCat 2.5 Preview FreeMeituanReleasedSep 25, 2026Capability indexNot tested yetBlended priceNot listedContext window1M
  11. MiMo V2.6 Flash UncensoredXiaomiReleasedSep 25, 2026Capability indexNot tested yetBlended price$0.75Context window1M
  12. Qwen 3.8 Max PrimeAlibabaReleasedSep 23, 2026Capability indexNot tested yetBlended price$6.00Context window1M

Capability

Rank models on the capability index or on any single test. Scores from different tests are never mixed.

Epoch Capabilities Index

One score built from many tests. Higher is better. The scale has no fixed top, so compare models with each other rather than reading it as a percentage.

Not on the index yet: MiMo-V2.6-Flash (Sep 22), MiMo-V2.6-Pro (Sep 22), Step 5 Preview (Sep 16), DeepSeek V4 Flash Vision Exp (Sep 10), Qwen3.8 Flash Next (Aug 27), Granite 4.2 8B (Aug 24). Epoch AI adds a model once enough of its tests have run; the results already published are on each model’s page.

The index has no zero point, so the axis starts near the lowest score shown. Gaps of a few points sit within its published range. Results published by Epoch AI (CC BY 4.0).

Full rankings

Score, price and time

Every model on the index against its price, release date and context window.

Capability index against price

Up and to the left is better: a higher score for less money. The dotted line links the models that no cheaper model beats.

  • Best score at each price
  • Above-median score, below-median price

Use the arrow keys to move between models. Press Enter to open the selected model. A table with the same data follows the chart.

Price is the maker’s own list price, or the middle price across hosts when the maker does not sell the model directly, blended as three input tokens for every output token. It is not the cost of any particular task. 183 of 274 models on the index are visible in this view.

Prices

What the most capable models cost to use, token by token.

Input, cached input and output prices

The fifteen most capable models with a list price, per 1M tokens. Output usually costs several times more than input, and cached input much less.

  1. Claude Opus 5.5$4.00 / $20.00$0.20 cached
  2. GPT-6 Astra$10.00 / $50.00$1.00 cached
  3. GPT-6.1 Sol$2.00 / $10.00$0.10 cached
  4. Claude Sonnet 5.5$2.00 / $10.00$0.20 cached
  5. Claude Fable 5.1$10.00 / $50.00$0.25 cached
  6. Claude Opus 5$5.00 / $25.00$0.50 cached
  7. GPT-6 Sol$2.00 / $10.00$0.20 cached
  8. GPT-5.5 Pro$30.00 / $180.00
  9. Claude Fable 5$10.00 / $50.00$1.00 cached
  10. GPT-5.6 Sol$4.00 / $20.00$0.40 cached
  11. GPT-5.6 Terra$2.00 / $12.00$0.20 cached
  12. GPT-5.5$5.00 / $30.00$0.50 cached
  13. GPT-5.4 Pro$30.00 / $180.00
  14. Claude Opus 4.8$5.00 / $25.00$0.50 cached
  15. Kimi K3$3.00 / $15.00$0.30 cached

Context and scale

How much text models take in at once, and how large their training runs were.

Largest context windows

Among models with a capability score. Tokens taken in at once.
  1. Gemini 2.0 Pro2.1M
  2. Grok 4.202M
  3. Grok 4 Fast2M
  4. GPT-6 Astra1.1M
  5. GPT-6.1 Sol1.1M
  6. GPT-6 Sol1.1M
  7. GPT-5.5 Pro1.1M
  8. GPT-5.6 Sol1.1M
  9. GPT-5.6 Terra1.1M
  10. GPT-5.51.1M
  11. GPT-5.4 Pro1.1M
  12. GPT-5.41.1M

Largest training runs

Estimated training compute, from Epoch AI's records. Log scale.
  1. GPT-6 Astra1.0 × 10²⁷ FLOP
  2. Grok 45.0 × 10²⁶ FLOP
  3. Grok 33.5 × 10²⁶ FLOP
  4. GPT-56.6 × 10²⁵ FLOP
  5. Claude 3.7 Sonnet3.4 × 10²⁵ FLOP
  6. Claude 3.5 Sonnet (October 2024)2.7 × 10²⁵ FLOP
  7. Claude 3.5 Sonnet2.7 × 10²⁵ FLOP
  8. Kimi K32.0 × 10²⁵ FLOP
  9. Qwen3 Max1.5 × 10²⁵ FLOP
  10. DeepSeek V4 Pro 08139.7 × 10²⁴ FLOP
  11. DeepSeek V4 Pro9.7 × 10²⁴ FLOP
  12. Llama 3.1-70B7.9 × 10²⁴ FLOP

Every model

All 1,996 models we track, highest capability score first. Search, filter or sort to find the right one.

Loading the model list…

Frequently asked questions

Answered from the data on this page.

Which is the most capable AI model?

Claude Opus 5.5 leads the Epoch Capabilities Index at 167.3, ahead of GPT-6 Astra at 166.4. Small gaps are within the index's published range.

What are the top AI models?

By capability index: Claude Opus 5.5, GPT-6 Astra, GPT-6.1 Sol, Claude Sonnet 5.5, Claude Fable 5.1.

Which is the cheapest AI model?

Among models with a capability score, Llama 3-8B has the lowest blended list price, $0.04 per 1M tokens. Smaller untested models can cost less.

Which AI model has the largest context window?

Llama 4 Scout 17B Instruct lists the largest context window, 10M tokens.

Which is the best open-weights AI model?

Kimi K3 scores highest among open-weights models on the capability index, at 157.4.

Which is the best reasoning model?

Claude Opus 5.5 scores highest among reasoning models on the capability index, at 167.3.

Which is the newest AI model?

Mistral Large 4 from Mistral AI is the newest release from a known maker, listed on Oct 6, 2026.

How are AI models compared on Artificials?

Capability comes from published tests: the Epoch Capabilities Index and Epoch AI evaluations, LiveBench and T2I-CoReBench. Prices, context windows and features come from the models.dev catalog. Artificials groups every listing of a model, links results to it and calculates the rankings and charts. Speed and latency are not measured yet.

How do I compare a specific model with others?

Open the model’s page from any chart or list for its scores, prices, providers and the models closest to it, or add up to three models to the side-by-side comparison.