GenAiHub

LLMs

Browse 435 LLMs

LLMs Guide

Each model is ranked by its published Intelligence Index score and shown alongside its context window, measured output speed, blended price per million tokens, release date and whether its weights are open. Scores come from published benchmark results, not from our own testing.

176–200 of 435

LLM leaderboard

Rank reflects Intelligence Index

Click a column to sort
LLM leaderboard ranked before pagination using published benchmark results.
RankPosition based on the Intelligence Index. Tied scores share the same rank.ModelModel name and publisher matched to the published benchmark result.Composite score across reasoning, knowledge, mathematics, science, and coding evaluations.Composite score for programming and software-engineering capability.Composite score for tool use, planning, and autonomous task completion.Maximum number of tokens the model can process in one request.Median output generation throughput, measured in tokens per second.Blended USD cost per 1 million tokens using a 3:1 input-to-output ratio.Public release month and year for this model.Whether downloadable model weights are open or access is closed.
176
OpenAI logo
o3-miniOpenAI
12.516.30.9200K204 tokens/s$1.93Jan 2025Closed source
177
OpenAI logo
o1-proOpenAI
12.4——200K—$262.5Mar 2025Closed source
178
Gr
Granite 4.2 8BIBM
12.422.43.7131.1K120 tokens/s$0.11Aug 2026Open source
179
OpenAI logo
gpt-oss-120bOpenAI
12.330.46.2131.1K208 tokens/s$0.26Aug 2025Open source
180
JT
JT-MINIChina Mobile
12.2——128K——Apr 2026Closed source
181
SpaceXAI logo
Grok 3SpaceXAI
12.1——1M—$8Feb 2025Closed source
182
Se
Seed-OSS-36B-InstructByteDance Seed
12.1——512K30 tokens/s$0.3Aug 2025Open source
183
Alibaba logo
Qwen3 235B 2507Alibaba
12.0——256K59 tokens/s$0.4Jul 2025Open source
184
Alibaba logo
Qwen3 Coder 480BAlibaba
11.9——262.1K57 tokens/s$3Jul 2025Open source
185
Alibaba logo
Qwen3 VL 32BAlibaba
11.9——256K90 tokens/s$0.28Oct 2025Open source
186
Li
Ling 3.0 TinyInclusionAI
11.926.57.1262.1K153 tokens/s$0Aug 2026Open source
187
Ma
Magistral Medium 1.2Mistral
11.821.3—128K——Sep 2025Closed source
188
So
Sonar Reasoning ProPerplexity
11.8——127K——Jan 2025Closed source
189
Hy
HyperNova 60B 2605Multiverse Computing
11.723.22.7131.1K353 tokens/s$0.07May 2026Open source
190
MiniMax logo
MiniMax M1 80kMiniMax
11.7——1M—$0.96Jun 2025Open source
191
Ne
Nemotron Cascade 2 30B A3BNVIDIA
11.725.3—1M——Mar 2026Open source
192
Me
Mercury 2Inception
11.531.14.0128K933 tokens/s$0.38Feb 2026Closed source
193
K2
K2 Think V2MBZUAI Institute of Foundation Models
11.521.0—262.1K——Dec 2025Open source
194
Lo
LongCat Flash LiteLongCat
11.5——256K——Jan 2026Open source
195
Mi
Mistral Small 4Mistral
11.526.61.4256K168 tokens/s$0.26Mar 2026Open source
196
DeepSeek logo
DeepSeek R1DeepSeek
11.424.61.1128K—$2.5Jan 2025Open source
197
OpenAI logo
o1-previewOpenAI
11.434.0—128K—$28.88Sep 2024Closed source
198
Hy
HyperCLOVA X SEED ThinkNaver
11.4——128K——Dec 2025Open source
199
Z AI logo
GLM-4.6VZ AI
11.2——128K63 tokens/s$0.45Dec 2025Open source
200
Alibaba logo
Qwen3 Next 80B A3BAlibaba
11.217.4—262.1K193 tokens/s$0.41Sep 2025Open source

Rank uses Intelligence Index. Sort any index, context, speed, pricing, release date, or license; missing data always appears last.

All benchmark leaderboards

Explore every published snapshot across coding, agentic search, reasoning, instruction following, long context, and safety evaluations.

Browse by category

31 of 31 benchmarks

Code & Software Engineering

SWE-bench Pro

A long-horizon repository benchmark built to test realistic software engineering work. Its scores need a task-quality and setup check before they support a coding-agent decision.

  1. 1
    Claude Opus 5.5 logo
    Claude Opus 5.589.9
  2. 2
    Claude Sonnet 5.5 logo
    Claude Sonnet 5.581.3
  3. 3
    Claude Fable 5.1 logo
    Claude Fable 5.181.2

78 models

Code & Software Engineering

SWE Multilingual

A multilingual software-engineering benchmark for real-world code issue resolution across multiple programming languages.

  1. 1
    Claude Opus 5.5 logo
    Claude Opus 5.593.9
  2. 2
    Claude Sonnet 5.5 logo
    Claude Sonnet 5.590.3
  3. 3
    Claude Opus 5 logo
    Claude Opus 589.5

48 models

Financial & Business Operations

AA Tau3 Banking

An independently evaluated Tau3 banking benchmark from Artificial Analysis.

  1. 1
    Grok 4.6 logo
    Grok 4.650.7
  2. 2
    Muse Spark 1.3 logo
    Muse Spark 1.350.5
  3. 3
    GLM-5.3 logo
    GLM-5.350.3

14 models

Task Planning & Knowledge Search

SkillsBench

How important are skills for agents?

  1. 1
    DeepSeek V4.1 Flash logo
    DeepSeek V4.1 Flash69.8
  2. 2
    Grok 4.5 logo
    Grok 4.566.0
  3. 3
    Gemini 3.7 Flash logo
    Gemini 3.7 Flash65.9

35 models

Task Planning & Knowledge Search

WideResearch

A broad research-agent benchmark for open-ended information gathering, synthesis, and answer construction across wide search spaces.

  1. 1
    Me
    Mercury 292.3
  2. 2
    Hy
    Hy4 preview83.9
  3. 3
    Qwen3.8 Max logo
    Qwen3.8 Max81.9

16 models

Task Planning & Knowledge Search

BrowseComp

A benchmark for web-browsing agents that must search, inspect sources, gather evidence, and return the correct answer to research-oriented questions.

  1. 1
    At
    Atria Dawn Preview92.5
  2. 2
    GPT-5.6 Sol logo
    GPT-5.6 Sol92.2
  3. 3
    GPT-6 Astra logo
    GPT-6 Astra91.5

47 models

Instruction Following & Long Context

MRCR v2 128K-256K

MRCR v2 slice focused on very long contexts at 128K-256K lengths.

  1. 1
    GPT-5.5 logo
    GPT-5.587.5
  2. 2
    Claude Opus 4.7 (Adaptive) logo
    Claude Opus 4.7 (Adaptive)59.2

2 models

Code & Software Engineering

SWE-bench Verified

Human-validated software engineering issues from real-world Python repositories.

  1. 1
    Claude Opus 4.7 (max) logo
    Claude Opus 4.7 (max)83.5
  2. 2
    GPT-5.5 (xhigh) logo
    GPT-5.5 (xhigh)80.6
  3. 3
    Gemini 3.5 Flash (high) logo
    Gemini 3.5 Flash (high)79.3

33 models

Knowledge

SimpleQA Verified

Factoid questions spanning politics, science, technology, art, sports, geography, and music.

  1. 1
    GPT-6 Astra (max) logo
    GPT-6 Astra (max)75.6
  2. 2
    GPT-6.1 Sol (max) logo
    GPT-6.1 Sol (max)73.9
  3. 3
    Gemini 3.1 Pro Preview (high) logo
    Gemini 3.1 Pro Preview (high)73.5

86 models

Reasoning

ARC-AGI-2

Abstract reasoning and pattern generalization on grid-based tasks.

  1. 1
    GPT-6 Astra (max) logo
    GPT-6 Astra (max)95.0
  2. 2
    GPT-6.1 Sol (max) logo
    GPT-6.1 Sol (max)94.2
  3. 3
    GPT-6 Astra (xhigh) logo
    GPT-6 Astra (xhigh)93.3

235 models

Math and science

FrontierMath Tier 4 (v2)

Exceptionally difficult research-level mathematics problems.

  1. 1
    GPT-6.1 Sol (max) logo
    GPT-6.1 Sol (max)100.0
  2. 2
    GPT-6 Astra (high) logo
    GPT-6 Astra (high)97.6
  3. 3
    GPT-6 Astra (xhigh) logo
    GPT-6 Astra (xhigh)97.6

70 models

Task Planning & Knowledge Search

MCP Atlas

Real-world multi-step tool-use evaluation through the Model Context Protocol.

  1. 1
    Muse Spark 1.1 logo
    Muse Spark 1.188.1
  2. 2
    Fa
    Fable 5.187.2
  3. 3
    claude-opus-5 (xhigh) logo
    claude-opus-5 (xhigh)85.8

34 models

Cybersecurity & Risk

CyberGym

Real-world cybersecurity evaluation of AI agents reproducing vulnerabilities with working proof-of-concept tests.

  1. 1
    Multi-model (DeepSeek-V4.1-Flash, Abliterated GLM-5.2) logo
    Multi-model (DeepSeek-V4.1-Flash, Abliterated GLM-5.2)99.2
  2. 2
    Multi-model (Creation Model, DeepSeek-V4-Pro, Qwen 3.8 Max) logo
    Multi-model (Creation Model, DeepSeek-V4-Pro, Qwen 3.8 Max)98.5
  3. 3
    Multi-model (DeepSeek-V4-Flash, OrcaCyber-Zero-1.0) logo
    Multi-model (DeepSeek-V4-Flash, OrcaCyber-Zero-1.0)98.0

77 models

Code & Software Engineering

Terminal-Bench 3.0

The official leaderboard for Terminal-Bench 3.0.

  1. 1
    Op
    Opus 542.7
  2. 2
    GPT-5.6 Sol logo
    GPT-5.6 Sol34.6
  3. 3
    Fa
    Fable 534.0

12 models

Instruction Following & Long Context

Humanity's Last Exam

Expert-level questions across mathematics, science, and humanities, as published by the CAIS AI Dashboard.

  1. 1
    Fa
    Fable 5.154.6
  2. 2
    GPT-6 Astra logo
    GPT-6 Astra53.6
  3. 3
    Fa
    Fable 552.7

60 models

Text Capabilities

Text Capabilities Index

Average of the text capability benchmarks published by the CAIS AI Dashboard.

  1. 1
    GPT-6 Astra logo
    GPT-6 Astra63.9
  2. 2
    Fa
    Fable 5.155.6
  3. 3
    Fa
    Fable 554.4

60 models

Vision Capabilities

Vision Capabilities Index

Average of the vision capability benchmarks published by the CAIS AI Dashboard.

  1. 1
    GPT-6 Astra logo
    GPT-6 Astra83.0
  2. 2
    Gemini 3.8 Flash logo
    Gemini 3.8 Flash69.9
  3. 3
    Fa
    Fable 5.168.9

48 models

Risk & Safety

Risk Index

Average of the risk and safety benchmarks published by the CAIS AI Dashboard; lower is better.

  1. 1
    Fa
    Fable 5.130.3
  2. 2
    Muse Spark 1.1 logo
    Muse Spark 1.132.5
  3. 3
    GPT-6 Astra logo
    GPT-6 Astra35.5

9 models

Automation

Automation

Remote Labor Index automation rates published by the CAIS AI Dashboard.

  1. 1
    Op
    Opus 5.521.3
  2. 2
    GPT-6 Astra logo
    GPT-6 Astra20.8
  3. 3
    Fa
    Fable 5.117.9

16 models

Reasoning

Intelligence Index

Composite Artificial Analysis index across mathematics, science, coding, and reasoning evaluations.

  1. 1
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)53.4
  2. 2
    Claude Fable 5.1 (xhigh with fallback) logo
    Claude Fable 5.1 (xhigh with fallback)53.2
  3. 3
    GPT-6 Astra (max) logo
    GPT-6 Astra (max)52.8

633 models

Code & Software Engineering

Coding Index

Artificial Analysis composite index for programming and software-engineering capability.

  1. 1
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)81.6
  2. 2
    Claude Fable 5.1 (xhigh with fallback) logo
    Claude Fable 5.1 (xhigh with fallback)80.7
  3. 3
    Claude Fable 5.1 (high with fallback) logo
    Claude Fable 5.1 (high with fallback)79.1

256 models

Task Planning & Knowledge Search

Agentic Index

Artificial Analysis composite index for tool use, planning, autonomy, and complex agentic workflows.

  1. 1
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)58.0
  2. 2
    Claude Fable 5.1 (xhigh with fallback) logo
    Claude Fable 5.1 (xhigh with fallback)57.1
  3. 3
    Claude Opus 5 (max) logo
    Claude Opus 5 (max)56.2

151 models

Knowledge

Omniscience Index

Artificial Analysis factual-knowledge and hallucination-resistance evaluation.

  1. 1
    GPT-6 Astra (high) logo
    GPT-6 Astra (high)43.7
  2. 2
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)43.5
  3. 3
    GPT-6 Astra (xhigh) logo
    GPT-6 Astra (xhigh)43.4

99 models

Knowledge

GPQA

Graduate-level, expert-written science questions evaluated by Artificial Analysis.

  1. 1
    GPT-6 Astra (xhigh) logo
    GPT-6 Astra (xhigh)96.3
  2. 2
    GPT-6 Astra (max) logo
    GPT-6 Astra (max)96.1
  3. 3
    Gemini 3.8 Flash (high) logo
    Gemini 3.8 Flash (high)95.3

612 models

Knowledge

Humanity's Last Exam

Broad expert-level reasoning and knowledge evaluation published by Artificial Analysis.

  1. 1
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)59.1
  2. 2
    Claude Fable 5.1 (xhigh with fallback) logo
    Claude Fable 5.1 (xhigh with fallback)58.7
  3. 3
    Claude Fable 5.1 (high with fallback) logo
    Claude Fable 5.1 (high with fallback)55.9

607 models

Instruction Following & Long Context

IFBench

Instruction-following benchmark results published by Artificial Analysis.

  1. 1
    Grok 4.3 (medium) logo
    Grok 4.3 (medium)83.3
  2. 2
    Grok 4.20 0309 logo
    Grok 4.20 030982.9
  3. 3
    MiniMax-M3 logo
    MiniMax-M382.9

450 models

Code & Software Engineering

SciCode

Scientific programming benchmark results published by Artificial Analysis.

  1. 1
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)63.1
  2. 2
    Claude Fable 5 (with fallback) logo
    Claude Fable 5 (with fallback)61.0
  3. 3
    Claude Fable 5.1 (xhigh with fallback) logo
    Claude Fable 5.1 (xhigh with fallback)60.9

167 models

Code & Software Engineering

Terminal-Bench 2.1

Agentic terminal and software-engineering results published by Artificial Analysis.

  1. 1
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)91.4
  2. 2
    Claude Fable 5.1 (xhigh with fallback) logo
    Claude Fable 5.1 (xhigh with fallback)91.0
  3. 3
    Claude Fable 5.1 (high with fallback) logo
    Claude Fable 5.1 (high with fallback)89.9

236 models

Code & Software Engineering

CritPt

Critical-points programming evaluation published by Artificial Analysis.

  1. 1
    GPT-5.6 Sol (max) logo
    GPT-5.6 Sol (max)32.3
  2. 2
    GPT-6 Astra (max) logo
    GPT-6 Astra (max)31.7
  3. 3
    GPT-6 Astra (xhigh) logo
    GPT-6 Astra (xhigh)31.4

520 models

Instruction Following & Long Context

Long Context Reasoning

Long-context reasoning results published by Artificial Analysis.

  1. 1
    Kimi K3 (max) logo
    Kimi K3 (max)88.7
  2. 2
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)85.3
  3. 3
    Claude Fable 5.1 (medium with fallback) logo
    Claude Fable 5.1 (medium with fallback)84.7

516 models

Task Planning & Knowledge Search

τ²-Bench

Tool-agent interaction benchmark results published by Artificial Analysis.

  1. 1
    JT
    JT-35B-Flash99.1
  2. 2
    GLM-5.2 (max) logo
    GLM-5.2 (max)99.1
  3. 3
    GLM-4.7-Flash logo
    GLM-4.7-Flash98.8

440 models

Scores are published snapshots and are only comparable within the same benchmark, version, metric, and evaluation configuration.