LLMs
Browse 435 LLMs
Each model is ranked by its published Intelligence Index score and shown alongside its context window, measured output speed, blended price per million tokens, release date and whether its weights are open. Scores come from published benchmark results, not from our own testing.
376–400 of 435
LLM leaderboard
Rank reflects Intelligence Index
| RankPosition based on the Intelligence Index. Tied scores share the same rank. | ModelModel name and publisher matched to the published benchmark result. | Composite score across reasoning, knowledge, mathematics, science, and coding evaluations. | Composite score for programming and software-engineering capability. | Composite score for tool use, planning, and autonomous task completion. | Maximum number of tokens the model can process in one request. | Median output generation throughput, measured in tokens per second. | Blended USD cost per 1 million tokens using a 3:1 input-to-output ratio. | Public release month and year for this model. | Whether downloadable model weights are open or access is closed. |
|---|---|---|---|---|---|---|---|---|---|
| 376 | Li | 5.5 | — | — | 131K | — | — | Sep 2025 | Open source |
| 377 | 5.5 | — | — | 128K | — | — | Jan 2025 | Open source | |
| 378 | 5.5 | 12.9 | — | 100K | — | — | Jul 2023 | Closed source | |
| 378 | 5.5 | — | — | 128K | — | — | May 2024 | Open source | |
| 380 | Mi | 5.5 | — | — | 32.8K | 137 tokens/s | $3 | Dec 2023 | Closed source |
| 381 | 5.5 | 10.7 | — | 4.1K | — | $0.75 | Nov 2022 | Closed source | |
| 382 | Mi | 5.5 | 9.7 | 0.6 | 256K | 89 tokens/s | $0.15 | Dec 2025 | Open source |
| 383 | 5.5 | — | — | 8.2K | — | $1.18 | Apr 2024 | Open source | |
| 384 | Ar | 5.4 | — | — | 4K | — | — | Apr 2024 | Open source |
| 384 | 5.4 | — | — | 33.8K | — | — | Nov 2023 | Open source | |
| 386 | LF | 5.4 | — | — | 32K | — | — | Sep 2024 | Closed source |
| 387 | 5.4 | — | — | 128K | 12 tokens/s | $0.35 | Sep 2024 | Open source | |
| 388 | PA | 5.4 | 4.6 | — | 8K | — | — | May 2023 | Closed source |
| 389 | 5.3 | — | — | 32.8K | — | — | Dec 2023 | Closed source | |
| 390 | 5.3 | — | — | 128K | — | — | Jun 2024 | Open source | |
| 391 | Sa | 5.3 | — | — | 32.8K | — | $0 | May 2025 | Open source |
| 392 | 5.3 | — | — | 4.1K | — | — | Nov 2023 | Open source | |
| 392 | 5.3 | — | — | 4.1K | — | — | Jul 2023 | Open source | |
| 394 | 5.3 | — | — | 4.1K | — | — | Jul 2023 | Open source | |
| 395 | Co | 5.3 | — | — | 128K | — | — | Mar 2024 | Open source |
| 396 | Op | 5.3 | — | — | 8.2K | — | — | Dec 2023 | Open source |
| 397 | DB | 5.3 | — | — | 32.8K | — | — | Mar 2024 | Open source |
| 398 | Ex | 5.3 | — | — | 64K | — | — | Jul 2025 | Open source |
| 399 | Ol | 5.2 | — | — | 65.5K | — | $0.13 | Nov 2025 | Open source |
| 400 | LF | 5.2 | — | — | 32K | — | — | Jan 2026 | Open source |
Rank uses Intelligence Index. Sort any index, context, speed, pricing, release date, or license; missing data always appears last. | |||||||||
All benchmark leaderboards
Explore every published snapshot across coding, agentic search, reasoning, instruction following, long context, and safety evaluations.
Browse by category
31 of 31 benchmarks
SWE-bench Pro
A long-horizon repository benchmark built to test realistic software engineering work. Its scores need a task-quality and setup check before they support a coding-agent decision.
- 1Claude Opus 5.589.9
- 2Claude Sonnet 5.581.3
- 3Claude Fable 5.181.2
78 models
SWE Multilingual
A multilingual software-engineering benchmark for real-world code issue resolution across multiple programming languages.
- 1Claude Opus 5.593.9
- 2Claude Sonnet 5.590.3
- 3Claude Opus 589.5
48 models
AA Tau3 Banking
An independently evaluated Tau3 banking benchmark from Artificial Analysis.
- 1Grok 4.650.7
- 2Muse Spark 1.350.5
- 3GLM-5.350.3
14 models
SkillsBench
How important are skills for agents?
- 1DeepSeek V4.1 Flash69.8
- 2Grok 4.566.0
- 3Gemini 3.7 Flash65.9
35 models
WideResearch
A broad research-agent benchmark for open-ended information gathering, synthesis, and answer construction across wide search spaces.
- 1MeMercury 292.3
- 2HyHy4 preview83.9
- 3Qwen3.8 Max81.9
16 models
BrowseComp
A benchmark for web-browsing agents that must search, inspect sources, gather evidence, and return the correct answer to research-oriented questions.
- 1AtAtria Dawn Preview92.5
- 2GPT-5.6 Sol92.2
- 3GPT-6 Astra91.5
47 models
MRCR v2 128K-256K
MRCR v2 slice focused on very long contexts at 128K-256K lengths.
- 1GPT-5.587.5
- 2Claude Opus 4.7 (Adaptive)59.2
2 models
SWE-bench Verified
Human-validated software engineering issues from real-world Python repositories.
- 1Claude Opus 4.7 (max)83.5
- 2GPT-5.5 (xhigh)80.6
- 3Gemini 3.5 Flash (high)79.3
33 models
SimpleQA Verified
Factoid questions spanning politics, science, technology, art, sports, geography, and music.
- 1GPT-6 Astra (max)75.6
- 2GPT-6.1 Sol (max)73.9
- 3Gemini 3.1 Pro Preview (high)73.5
86 models
ARC-AGI-2
Abstract reasoning and pattern generalization on grid-based tasks.
- 1GPT-6 Astra (max)95.0
- 2GPT-6.1 Sol (max)94.2
- 3GPT-6 Astra (xhigh)93.3
235 models
FrontierMath Tier 4 (v2)
Exceptionally difficult research-level mathematics problems.
- 1GPT-6.1 Sol (max)100.0
- 2GPT-6 Astra (high)97.6
- 3GPT-6 Astra (xhigh)97.6
70 models
MCP Atlas
Real-world multi-step tool-use evaluation through the Model Context Protocol.
- 1Muse Spark 1.188.1
- 2FaFable 5.187.2
- 3claude-opus-5 (xhigh)85.8
34 models
CyberGym
Real-world cybersecurity evaluation of AI agents reproducing vulnerabilities with working proof-of-concept tests.
- 1Multi-model (DeepSeek-V4.1-Flash, Abliterated GLM-5.2)99.2
- 2Multi-model (Creation Model, DeepSeek-V4-Pro, Qwen 3.8 Max)98.5
- 3Multi-model (DeepSeek-V4-Flash, OrcaCyber-Zero-1.0)98.0
77 models
Terminal-Bench 3.0
The official leaderboard for Terminal-Bench 3.0.
- 1OpOpus 542.7
- 2GPT-5.6 Sol34.6
- 3FaFable 534.0
12 models
Humanity's Last Exam
Expert-level questions across mathematics, science, and humanities, as published by the CAIS AI Dashboard.
- 1FaFable 5.154.6
- 2GPT-6 Astra53.6
- 3FaFable 552.7
60 models
Text Capabilities Index
Average of the text capability benchmarks published by the CAIS AI Dashboard.
- 1GPT-6 Astra63.9
- 2FaFable 5.155.6
- 3FaFable 554.4
60 models
Vision Capabilities Index
Average of the vision capability benchmarks published by the CAIS AI Dashboard.
- 1GPT-6 Astra83.0
- 2Gemini 3.8 Flash69.9
- 3FaFable 5.168.9
48 models
Risk Index
Average of the risk and safety benchmarks published by the CAIS AI Dashboard; lower is better.
- 1FaFable 5.130.3
- 2Muse Spark 1.132.5
- 3GPT-6 Astra35.5
9 models
Automation
Remote Labor Index automation rates published by the CAIS AI Dashboard.
- 1OpOpus 5.521.3
- 2GPT-6 Astra20.8
- 3FaFable 5.117.9
16 models
Intelligence Index
Composite Artificial Analysis index across mathematics, science, coding, and reasoning evaluations.
- 1Claude Fable 5.1 (max with fallback)53.4
- 2Claude Fable 5.1 (xhigh with fallback)53.2
- 3GPT-6 Astra (max)52.8
633 models
Coding Index
Artificial Analysis composite index for programming and software-engineering capability.
- 1Claude Fable 5.1 (max with fallback)81.6
- 2Claude Fable 5.1 (xhigh with fallback)80.7
- 3Claude Fable 5.1 (high with fallback)79.1
256 models
Agentic Index
Artificial Analysis composite index for tool use, planning, autonomy, and complex agentic workflows.
- 1Claude Fable 5.1 (max with fallback)58.0
- 2Claude Fable 5.1 (xhigh with fallback)57.1
- 3Claude Opus 5 (max)56.2
151 models
Omniscience Index
Artificial Analysis factual-knowledge and hallucination-resistance evaluation.
- 1GPT-6 Astra (high)43.7
- 2Claude Fable 5.1 (max with fallback)43.5
- 3GPT-6 Astra (xhigh)43.4
99 models
GPQA
Graduate-level, expert-written science questions evaluated by Artificial Analysis.
- 1GPT-6 Astra (xhigh)96.3
- 2GPT-6 Astra (max)96.1
- 3Gemini 3.8 Flash (high)95.3
612 models
Humanity's Last Exam
Broad expert-level reasoning and knowledge evaluation published by Artificial Analysis.
- 1Claude Fable 5.1 (max with fallback)59.1
- 2Claude Fable 5.1 (xhigh with fallback)58.7
- 3Claude Fable 5.1 (high with fallback)55.9
607 models
IFBench
Instruction-following benchmark results published by Artificial Analysis.
- 1Grok 4.3 (medium)83.3
- 2Grok 4.20 030982.9
- 3MiniMax-M382.9
450 models
SciCode
Scientific programming benchmark results published by Artificial Analysis.
- 1Claude Fable 5.1 (max with fallback)63.1
- 2Claude Fable 5 (with fallback)61.0
- 3Claude Fable 5.1 (xhigh with fallback)60.9
167 models
Terminal-Bench 2.1
Agentic terminal and software-engineering results published by Artificial Analysis.
- 1Claude Fable 5.1 (max with fallback)91.4
- 2Claude Fable 5.1 (xhigh with fallback)91.0
- 3Claude Fable 5.1 (high with fallback)89.9
236 models
CritPt
Critical-points programming evaluation published by Artificial Analysis.
- 1GPT-5.6 Sol (max)32.3
- 2GPT-6 Astra (max)31.7
- 3GPT-6 Astra (xhigh)31.4
520 models
Long Context Reasoning
Long-context reasoning results published by Artificial Analysis.
- 1Kimi K3 (max)88.7
- 2Claude Fable 5.1 (max with fallback)85.3
- 3Claude Fable 5.1 (medium with fallback)84.7
516 models
τ²-Bench
Tool-agent interaction benchmark results published by Artificial Analysis.
- 1JTJT-35B-Flash99.1
- 2GLM-5.2 (max)99.1
- 3GLM-4.7-Flash98.8
440 models
Scores are published snapshots and are only comparable within the same benchmark, version, metric, and evaluation configuration.