GenAiHub

LLMs

Browse 435 LLMs

LLMs Guide

Each model is ranked by its published Intelligence Index score and shown alongside its context window, measured output speed, blended price per million tokens, release date and whether its weights are open. Scores come from published benchmark results, not from our own testing.

151–175 of 435

LLM leaderboard

Rank reflects Intelligence Index

Click a column to sort
LLM leaderboard ranked before pagination using published benchmark results.
RankPosition based on the Intelligence Index. Tied scores share the same rank.ModelModel name and publisher matched to the published benchmark result.Composite score across reasoning, knowledge, mathematics, science, and coding evaluations.Composite score for programming and software-engineering capability.Composite score for tool use, planning, and autonomous task completion.Maximum number of tokens the model can process in one request.Median output generation throughput, measured in tokens per second.Blended USD cost per 1 million tokens using a 3:1 input-to-output ratio.Public release month and year for this model.Whether downloadable model weights are open or access is closed.
151
Mi
MiniCPM5-2Bopenbmb
14.314.59.0131.1K——Sep 2026Open source
152
ER
ERNIE 5.0 Thinking PreviewBaidu
14.3——128K——Nov 2025Closed source
153
Ge
Gemma 4 12BGoogle
14.231.0—256K108 tokens/s$0.15Jun 2026Open source
154
No
Nova 2.0 Pro PreviewAmazon
14.234.0—256K123 tokens/s$3.44Nov 2025Closed source
155
SpaceXAI logo
Grok Code Fast 1SpaceXAI
14.1——256K——Aug 2025Closed source
156
Co
Command A+Cohere
13.927.83.6192K233 tokens/s$0May 2026Open source
157
Ap
Apriel-v1.5-15B-ThinkerServiceNow
13.8——128K—$0Sep 2025Open source
158
DeepSeek logo
DeepSeek V3.1DeepSeek
13.7——128K—$0.85Aug 2025Open source
159
Alibaba logo
Qwen3.5 9BAlibaba
13.728.7—262.1K83 tokens/s$0.15Mar 2026Open source
160
No
Nova 2.0 OmniAmazon
13.6——1M—$0.85Nov 2025Closed source
161
Ne
Nemotron 3.5 LightningNVIDIA
13.626.86.11M284 tokens/s$0.1Aug 2026Open source
162
Ne
Nemotron 3 SuperNVIDIA
13.637.74.11M99 tokens/s$0.31Mar 2026Open source
163
Alibaba logo
Qwen3 VL 235B A22BAlibaba
13.4——262.1K55 tokens/s$1.3Sep 2025Open source
164
Ap
Apriel-v1.6-15B-ThinkerServiceNow
13.4——128K—$0Nov 2025Open source
165
No
Nova 2.0 LiteAmazon
13.423.0—1M161 tokens/s$0.85Oct 2025Closed source
166
EX
EXAONE 4.5 33BLG AI Research
13.223.6—262.1K——Apr 2026Open source
167
Alibaba logo
Qwen3.5 4BAlibaba
13.122.6—262.1K27 tokens/s$0.06Mar 2026Open source
168
DeepSeek logo
DeepSeek R1 0528DeepSeek
13.1——128K—$1.76May 2025Open source
169
OpenAI logo
GPT-5 nanoOpenAI
13.0——400K186 tokens/s$0.14Aug 2025Closed source
170
No
North Mini CodeCohere
12.836.51.1256K44 tokens/s$0Jun 2026Open source
171
Z AI logo
GLM-4.5Z AI
12.8——128K——Jul 2025Open source
172
Kimi logo
Kimi K2Kimi
12.7——128K40 tokens/s$1Jul 2025Open source
173
Alibaba logo
Qwen3 235B A22B 2507Alibaba
12.722.11.3256K58 tokens/s$0.75Jul 2025Open source
174
OpenAI logo
GPT-4.1OpenAI
12.7——1M161 tokens/s$3.5Apr 2025Closed source
175
Alibaba logo
Qwen3.5 Omni FlashAlibaba
12.5——256K248 tokens/s$0.28Mar 2026Closed source

Rank uses Intelligence Index. Sort any index, context, speed, pricing, release date, or license; missing data always appears last.

All benchmark leaderboards

Explore every published snapshot across coding, agentic search, reasoning, instruction following, long context, and safety evaluations.

Browse by category

31 of 31 benchmarks

Code & Software Engineering

SWE-bench Pro

A long-horizon repository benchmark built to test realistic software engineering work. Its scores need a task-quality and setup check before they support a coding-agent decision.

  1. 1
    Claude Opus 5.5 logo
    Claude Opus 5.589.9
  2. 2
    Claude Sonnet 5.5 logo
    Claude Sonnet 5.581.3
  3. 3
    Claude Fable 5.1 logo
    Claude Fable 5.181.2

78 models

Code & Software Engineering

SWE Multilingual

A multilingual software-engineering benchmark for real-world code issue resolution across multiple programming languages.

  1. 1
    Claude Opus 5.5 logo
    Claude Opus 5.593.9
  2. 2
    Claude Sonnet 5.5 logo
    Claude Sonnet 5.590.3
  3. 3
    Claude Opus 5 logo
    Claude Opus 589.5

48 models

Financial & Business Operations

AA Tau3 Banking

An independently evaluated Tau3 banking benchmark from Artificial Analysis.

  1. 1
    Grok 4.6 logo
    Grok 4.650.7
  2. 2
    Muse Spark 1.3 logo
    Muse Spark 1.350.5
  3. 3
    GLM-5.3 logo
    GLM-5.350.3

14 models

Task Planning & Knowledge Search

SkillsBench

How important are skills for agents?

  1. 1
    DeepSeek V4.1 Flash logo
    DeepSeek V4.1 Flash69.8
  2. 2
    Grok 4.5 logo
    Grok 4.566.0
  3. 3
    Gemini 3.7 Flash logo
    Gemini 3.7 Flash65.9

35 models

Task Planning & Knowledge Search

WideResearch

A broad research-agent benchmark for open-ended information gathering, synthesis, and answer construction across wide search spaces.

  1. 1
    Me
    Mercury 292.3
  2. 2
    Hy
    Hy4 preview83.9
  3. 3
    Qwen3.8 Max logo
    Qwen3.8 Max81.9

16 models

Task Planning & Knowledge Search

BrowseComp

A benchmark for web-browsing agents that must search, inspect sources, gather evidence, and return the correct answer to research-oriented questions.

  1. 1
    At
    Atria Dawn Preview92.5
  2. 2
    GPT-5.6 Sol logo
    GPT-5.6 Sol92.2
  3. 3
    GPT-6 Astra logo
    GPT-6 Astra91.5

47 models

Instruction Following & Long Context

MRCR v2 128K-256K

MRCR v2 slice focused on very long contexts at 128K-256K lengths.

  1. 1
    GPT-5.5 logo
    GPT-5.587.5
  2. 2
    Claude Opus 4.7 (Adaptive) logo
    Claude Opus 4.7 (Adaptive)59.2

2 models

Code & Software Engineering

SWE-bench Verified

Human-validated software engineering issues from real-world Python repositories.

  1. 1
    Claude Opus 4.7 (max) logo
    Claude Opus 4.7 (max)83.5
  2. 2
    GPT-5.5 (xhigh) logo
    GPT-5.5 (xhigh)80.6
  3. 3
    Gemini 3.5 Flash (high) logo
    Gemini 3.5 Flash (high)79.3

33 models

Knowledge

SimpleQA Verified

Factoid questions spanning politics, science, technology, art, sports, geography, and music.

  1. 1
    GPT-6 Astra (max) logo
    GPT-6 Astra (max)75.6
  2. 2
    GPT-6.1 Sol (max) logo
    GPT-6.1 Sol (max)73.9
  3. 3
    Gemini 3.1 Pro Preview (high) logo
    Gemini 3.1 Pro Preview (high)73.5

86 models

Reasoning

ARC-AGI-2

Abstract reasoning and pattern generalization on grid-based tasks.

  1. 1
    GPT-6 Astra (max) logo
    GPT-6 Astra (max)95.0
  2. 2
    GPT-6.1 Sol (max) logo
    GPT-6.1 Sol (max)94.2
  3. 3
    GPT-6 Astra (xhigh) logo
    GPT-6 Astra (xhigh)93.3

235 models

Math and science

FrontierMath Tier 4 (v2)

Exceptionally difficult research-level mathematics problems.

  1. 1
    GPT-6.1 Sol (max) logo
    GPT-6.1 Sol (max)100.0
  2. 2
    GPT-6 Astra (high) logo
    GPT-6 Astra (high)97.6
  3. 3
    GPT-6 Astra (xhigh) logo
    GPT-6 Astra (xhigh)97.6

70 models

Task Planning & Knowledge Search

MCP Atlas

Real-world multi-step tool-use evaluation through the Model Context Protocol.

  1. 1
    Muse Spark 1.1 logo
    Muse Spark 1.188.1
  2. 2
    Fa
    Fable 5.187.2
  3. 3
    claude-opus-5 (xhigh) logo
    claude-opus-5 (xhigh)85.8

34 models

Cybersecurity & Risk

CyberGym

Real-world cybersecurity evaluation of AI agents reproducing vulnerabilities with working proof-of-concept tests.

  1. 1
    Multi-model (DeepSeek-V4.1-Flash, Abliterated GLM-5.2) logo
    Multi-model (DeepSeek-V4.1-Flash, Abliterated GLM-5.2)99.2
  2. 2
    Multi-model (Creation Model, DeepSeek-V4-Pro, Qwen 3.8 Max) logo
    Multi-model (Creation Model, DeepSeek-V4-Pro, Qwen 3.8 Max)98.5
  3. 3
    Multi-model (DeepSeek-V4-Flash, OrcaCyber-Zero-1.0) logo
    Multi-model (DeepSeek-V4-Flash, OrcaCyber-Zero-1.0)98.0

77 models

Code & Software Engineering

Terminal-Bench 3.0

The official leaderboard for Terminal-Bench 3.0.

  1. 1
    Op
    Opus 542.7
  2. 2
    GPT-5.6 Sol logo
    GPT-5.6 Sol34.6
  3. 3
    Fa
    Fable 534.0

12 models

Instruction Following & Long Context

Humanity's Last Exam

Expert-level questions across mathematics, science, and humanities, as published by the CAIS AI Dashboard.

  1. 1
    Fa
    Fable 5.154.6
  2. 2
    GPT-6 Astra logo
    GPT-6 Astra53.6
  3. 3
    Fa
    Fable 552.7

60 models

Text Capabilities

Text Capabilities Index

Average of the text capability benchmarks published by the CAIS AI Dashboard.

  1. 1
    GPT-6 Astra logo
    GPT-6 Astra63.9
  2. 2
    Fa
    Fable 5.155.6
  3. 3
    Fa
    Fable 554.4

60 models

Vision Capabilities

Vision Capabilities Index

Average of the vision capability benchmarks published by the CAIS AI Dashboard.

  1. 1
    GPT-6 Astra logo
    GPT-6 Astra83.0
  2. 2
    Gemini 3.8 Flash logo
    Gemini 3.8 Flash69.9
  3. 3
    Fa
    Fable 5.168.9

48 models

Risk & Safety

Risk Index

Average of the risk and safety benchmarks published by the CAIS AI Dashboard; lower is better.

  1. 1
    Fa
    Fable 5.130.3
  2. 2
    Muse Spark 1.1 logo
    Muse Spark 1.132.5
  3. 3
    GPT-6 Astra logo
    GPT-6 Astra35.5

9 models

Automation

Automation

Remote Labor Index automation rates published by the CAIS AI Dashboard.

  1. 1
    Op
    Opus 5.521.3
  2. 2
    GPT-6 Astra logo
    GPT-6 Astra20.8
  3. 3
    Fa
    Fable 5.117.9

16 models

Reasoning

Intelligence Index

Composite Artificial Analysis index across mathematics, science, coding, and reasoning evaluations.

  1. 1
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)53.4
  2. 2
    Claude Fable 5.1 (xhigh with fallback) logo
    Claude Fable 5.1 (xhigh with fallback)53.2
  3. 3
    GPT-6 Astra (max) logo
    GPT-6 Astra (max)52.8

633 models

Code & Software Engineering

Coding Index

Artificial Analysis composite index for programming and software-engineering capability.

  1. 1
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)81.6
  2. 2
    Claude Fable 5.1 (xhigh with fallback) logo
    Claude Fable 5.1 (xhigh with fallback)80.7
  3. 3
    Claude Fable 5.1 (high with fallback) logo
    Claude Fable 5.1 (high with fallback)79.1

256 models

Task Planning & Knowledge Search

Agentic Index

Artificial Analysis composite index for tool use, planning, autonomy, and complex agentic workflows.

  1. 1
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)58.0
  2. 2
    Claude Fable 5.1 (xhigh with fallback) logo
    Claude Fable 5.1 (xhigh with fallback)57.1
  3. 3
    Claude Opus 5 (max) logo
    Claude Opus 5 (max)56.2

151 models

Knowledge

Omniscience Index

Artificial Analysis factual-knowledge and hallucination-resistance evaluation.

  1. 1
    GPT-6 Astra (high) logo
    GPT-6 Astra (high)43.7
  2. 2
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)43.5
  3. 3
    GPT-6 Astra (xhigh) logo
    GPT-6 Astra (xhigh)43.4

99 models

Knowledge

GPQA

Graduate-level, expert-written science questions evaluated by Artificial Analysis.

  1. 1
    GPT-6 Astra (xhigh) logo
    GPT-6 Astra (xhigh)96.3
  2. 2
    GPT-6 Astra (max) logo
    GPT-6 Astra (max)96.1
  3. 3
    Gemini 3.8 Flash (high) logo
    Gemini 3.8 Flash (high)95.3

612 models

Knowledge

Humanity's Last Exam

Broad expert-level reasoning and knowledge evaluation published by Artificial Analysis.

  1. 1
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)59.1
  2. 2
    Claude Fable 5.1 (xhigh with fallback) logo
    Claude Fable 5.1 (xhigh with fallback)58.7
  3. 3
    Claude Fable 5.1 (high with fallback) logo
    Claude Fable 5.1 (high with fallback)55.9

607 models

Instruction Following & Long Context

IFBench

Instruction-following benchmark results published by Artificial Analysis.

  1. 1
    Grok 4.3 (medium) logo
    Grok 4.3 (medium)83.3
  2. 2
    Grok 4.20 0309 logo
    Grok 4.20 030982.9
  3. 3
    MiniMax-M3 logo
    MiniMax-M382.9

450 models

Code & Software Engineering

SciCode

Scientific programming benchmark results published by Artificial Analysis.

  1. 1
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)63.1
  2. 2
    Claude Fable 5 (with fallback) logo
    Claude Fable 5 (with fallback)61.0
  3. 3
    Claude Fable 5.1 (xhigh with fallback) logo
    Claude Fable 5.1 (xhigh with fallback)60.9

167 models

Code & Software Engineering

Terminal-Bench 2.1

Agentic terminal and software-engineering results published by Artificial Analysis.

  1. 1
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)91.4
  2. 2
    Claude Fable 5.1 (xhigh with fallback) logo
    Claude Fable 5.1 (xhigh with fallback)91.0
  3. 3
    Claude Fable 5.1 (high with fallback) logo
    Claude Fable 5.1 (high with fallback)89.9

236 models

Code & Software Engineering

CritPt

Critical-points programming evaluation published by Artificial Analysis.

  1. 1
    GPT-5.6 Sol (max) logo
    GPT-5.6 Sol (max)32.3
  2. 2
    GPT-6 Astra (max) logo
    GPT-6 Astra (max)31.7
  3. 3
    GPT-6 Astra (xhigh) logo
    GPT-6 Astra (xhigh)31.4

520 models

Instruction Following & Long Context

Long Context Reasoning

Long-context reasoning results published by Artificial Analysis.

  1. 1
    Kimi K3 (max) logo
    Kimi K3 (max)88.7
  2. 2
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)85.3
  3. 3
    Claude Fable 5.1 (medium with fallback) logo
    Claude Fable 5.1 (medium with fallback)84.7

516 models

Task Planning & Knowledge Search

τ²-Bench

Tool-agent interaction benchmark results published by Artificial Analysis.

  1. 1
    JT
    JT-35B-Flash99.1
  2. 2
    GLM-5.2 (max) logo
    GLM-5.2 (max)99.1
  3. 3
    GLM-4.7-Flash logo
    GLM-4.7-Flash98.8

440 models

Scores are published snapshots and are only comparable within the same benchmark, version, metric, and evaluation configuration.