GenAiHub

LLMs

Browse 435 LLMs

LLMs Guide

Each model is ranked by its published Intelligence Index score and shown alongside its context window, measured output speed, blended price per million tokens, release date and whether its weights are open. Scores come from published benchmark results, not from our own testing.

126–150 of 435

LLM leaderboard

Rank reflects Intelligence Index

Click a column to sort
LLM leaderboard ranked before pagination using published benchmark results.
RankPosition based on the Intelligence Index. Tied scores share the same rank.ModelModel name and publisher matched to the published benchmark result.Composite score across reasoning, knowledge, mathematics, science, and coding evaluations.Composite score for programming and software-engineering capability.Composite score for tool use, planning, and autonomous task completion.Maximum number of tokens the model can process in one request.Median output generation throughput, measured in tokens per second.Blended USD cost per 1 million tokens using a 3:1 input-to-output ratio.Public release month and year for this model.Whether downloadable model weights are open or access is closed.
126
Alibaba logo
Qwen3.5 122B A10BAlibaba
17.745.79.6262.1K144 tokens/s$1.1Feb 2026Open source
127
Anthropic logo
Claude 3.7 SonnetAnthropic
17.736.4—200K——Feb 2025Closed source
128
Anthropic logo
Claude 4.5 HaikuAnthropic
17.643.910.3200K83 tokens/s$2Oct 2025Closed source
129
Ri
Ring-2.6-1TInclusionAI
17.342.812.9262.1K121 tokens/s$0.85May 2026Open source
130
Li
Ling-2.6-1TInclusionAI
17.0——262.1K—$0.85Apr 2026Open source
131
StepFun logo
Step 3.5 Flash 2603StepFun
17.0——256K218 tokens/s$0.15Apr 2026Closed source
132
Do
Doubao Seed CodeByteDance Seed
16.9——256K——Nov 2025Closed source
133
Ge
Gemma 4 26B A4BGoogle
16.739.3—256K—$0.18Apr 2026Open source
134
Google logo
Gemini 2.5 ProGoogle
16.746.73.51M127 tokens/s$3.44Jun 2025Closed source
135
OpenAI logo
o4-miniOpenAI
16.7——200K157 tokens/s$1.93Apr 2025Closed source
136
StepFun logo
Step 3.5 FlashStepFun
16.6——256K226 tokens/s$0.15Feb 2026Open source
137
DeepSeek logo
DeepSeek V3.2 ExpDeepSeek
16.6——128K—$0.32Sep 2025Open source
138
Google logo
Gemini 3.1 Flash-LiteGoogle
16.034.73.21M315 tokens/s$0.56Mar 2026Closed source
139
Alibaba logo
Qwen3 MaxAlibaba
15.6——262.1K53 tokens/s$2.4Sep 2025Closed source
140
Google logo
Gemini 2.5 FlashGoogle
15.5——1M——Apr 2025Closed source
141
Ge
Gemma 4 31BGoogle
15.443.46.7256K35 tokens/s$0Apr 2026Open source
142
DeepSeek logo
DeepSeek V3.1 TerminusDeepSeek
15.443.58.9128K—$1.91Sep 2025Open source
143
Kimi logo
Kimi K2 0905Kimi
15.3——256K40 tokens/s$1.08Sep 2025Open source
144
OpenAI logo
o1OpenAI
15.239.7—200K—$26.25Dec 2024Closed source
145
Mi
Mistral Medium 3.5Mistral
14.946.99.4256K148 tokens/s$3Apr 2026Open source
146
Z AI logo
GLM-4.7-FlashZ AI
14.9——200K103 tokens/s$0.15Jan 2026Open source
147
Gr
Granite 4.2 30BIBM
14.829.9—131.1K76 tokens/s$0.28Aug 2026Open source
148
SpaceXAI logo
Grok 3 mini ReasoningSpaceXAI
14.6——1M77 tokens/s$0.35Feb 2025Closed source
149
DeepSeek logo
DeepSeek V3.2 SpecialeDeepSeek
14.5——128K——Dec 2025Open source
150
K-
K-EXAONELG AI Research
14.432.1—256K——Dec 2025Open source

Rank uses Intelligence Index. Sort any index, context, speed, pricing, release date, or license; missing data always appears last.

All benchmark leaderboards

Explore every published snapshot across coding, agentic search, reasoning, instruction following, long context, and safety evaluations.

Browse by category

31 of 31 benchmarks

Code & Software Engineering

SWE-bench Pro

A long-horizon repository benchmark built to test realistic software engineering work. Its scores need a task-quality and setup check before they support a coding-agent decision.

  1. 1
    Claude Opus 5.5 logo
    Claude Opus 5.589.9
  2. 2
    Claude Sonnet 5.5 logo
    Claude Sonnet 5.581.3
  3. 3
    Claude Fable 5.1 logo
    Claude Fable 5.181.2

78 models

Code & Software Engineering

SWE Multilingual

A multilingual software-engineering benchmark for real-world code issue resolution across multiple programming languages.

  1. 1
    Claude Opus 5.5 logo
    Claude Opus 5.593.9
  2. 2
    Claude Sonnet 5.5 logo
    Claude Sonnet 5.590.3
  3. 3
    Claude Opus 5 logo
    Claude Opus 589.5

48 models

Financial & Business Operations

AA Tau3 Banking

An independently evaluated Tau3 banking benchmark from Artificial Analysis.

  1. 1
    Grok 4.6 logo
    Grok 4.650.7
  2. 2
    Muse Spark 1.3 logo
    Muse Spark 1.350.5
  3. 3
    GLM-5.3 logo
    GLM-5.350.3

14 models

Task Planning & Knowledge Search

SkillsBench

How important are skills for agents?

  1. 1
    DeepSeek V4.1 Flash logo
    DeepSeek V4.1 Flash69.8
  2. 2
    Grok 4.5 logo
    Grok 4.566.0
  3. 3
    Gemini 3.7 Flash logo
    Gemini 3.7 Flash65.9

35 models

Task Planning & Knowledge Search

WideResearch

A broad research-agent benchmark for open-ended information gathering, synthesis, and answer construction across wide search spaces.

  1. 1
    Me
    Mercury 292.3
  2. 2
    Hy
    Hy4 preview83.9
  3. 3
    Qwen3.8 Max logo
    Qwen3.8 Max81.9

16 models

Task Planning & Knowledge Search

BrowseComp

A benchmark for web-browsing agents that must search, inspect sources, gather evidence, and return the correct answer to research-oriented questions.

  1. 1
    At
    Atria Dawn Preview92.5
  2. 2
    GPT-5.6 Sol logo
    GPT-5.6 Sol92.2
  3. 3
    GPT-6 Astra logo
    GPT-6 Astra91.5

47 models

Instruction Following & Long Context

MRCR v2 128K-256K

MRCR v2 slice focused on very long contexts at 128K-256K lengths.

  1. 1
    GPT-5.5 logo
    GPT-5.587.5
  2. 2
    Claude Opus 4.7 (Adaptive) logo
    Claude Opus 4.7 (Adaptive)59.2

2 models

Code & Software Engineering

SWE-bench Verified

Human-validated software engineering issues from real-world Python repositories.

  1. 1
    Claude Opus 4.7 (max) logo
    Claude Opus 4.7 (max)83.5
  2. 2
    GPT-5.5 (xhigh) logo
    GPT-5.5 (xhigh)80.6
  3. 3
    Gemini 3.5 Flash (high) logo
    Gemini 3.5 Flash (high)79.3

33 models

Knowledge

SimpleQA Verified

Factoid questions spanning politics, science, technology, art, sports, geography, and music.

  1. 1
    GPT-6 Astra (max) logo
    GPT-6 Astra (max)75.6
  2. 2
    GPT-6.1 Sol (max) logo
    GPT-6.1 Sol (max)73.9
  3. 3
    Gemini 3.1 Pro Preview (high) logo
    Gemini 3.1 Pro Preview (high)73.5

86 models

Reasoning

ARC-AGI-2

Abstract reasoning and pattern generalization on grid-based tasks.

  1. 1
    GPT-6 Astra (max) logo
    GPT-6 Astra (max)95.0
  2. 2
    GPT-6.1 Sol (max) logo
    GPT-6.1 Sol (max)94.2
  3. 3
    GPT-6 Astra (xhigh) logo
    GPT-6 Astra (xhigh)93.3

235 models

Math and science

FrontierMath Tier 4 (v2)

Exceptionally difficult research-level mathematics problems.

  1. 1
    GPT-6.1 Sol (max) logo
    GPT-6.1 Sol (max)100.0
  2. 2
    GPT-6 Astra (high) logo
    GPT-6 Astra (high)97.6
  3. 3
    GPT-6 Astra (xhigh) logo
    GPT-6 Astra (xhigh)97.6

70 models

Task Planning & Knowledge Search

MCP Atlas

Real-world multi-step tool-use evaluation through the Model Context Protocol.

  1. 1
    Muse Spark 1.1 logo
    Muse Spark 1.188.1
  2. 2
    Fa
    Fable 5.187.2
  3. 3
    claude-opus-5 (xhigh) logo
    claude-opus-5 (xhigh)85.8

34 models

Cybersecurity & Risk

CyberGym

Real-world cybersecurity evaluation of AI agents reproducing vulnerabilities with working proof-of-concept tests.

  1. 1
    Multi-model (DeepSeek-V4.1-Flash, Abliterated GLM-5.2) logo
    Multi-model (DeepSeek-V4.1-Flash, Abliterated GLM-5.2)99.2
  2. 2
    Multi-model (Creation Model, DeepSeek-V4-Pro, Qwen 3.8 Max) logo
    Multi-model (Creation Model, DeepSeek-V4-Pro, Qwen 3.8 Max)98.5
  3. 3
    Multi-model (DeepSeek-V4-Flash, OrcaCyber-Zero-1.0) logo
    Multi-model (DeepSeek-V4-Flash, OrcaCyber-Zero-1.0)98.0

77 models

Code & Software Engineering

Terminal-Bench 3.0

The official leaderboard for Terminal-Bench 3.0.

  1. 1
    Op
    Opus 542.7
  2. 2
    GPT-5.6 Sol logo
    GPT-5.6 Sol34.6
  3. 3
    Fa
    Fable 534.0

12 models

Instruction Following & Long Context

Humanity's Last Exam

Expert-level questions across mathematics, science, and humanities, as published by the CAIS AI Dashboard.

  1. 1
    Fa
    Fable 5.154.6
  2. 2
    GPT-6 Astra logo
    GPT-6 Astra53.6
  3. 3
    Fa
    Fable 552.7

60 models

Text Capabilities

Text Capabilities Index

Average of the text capability benchmarks published by the CAIS AI Dashboard.

  1. 1
    GPT-6 Astra logo
    GPT-6 Astra63.9
  2. 2
    Fa
    Fable 5.155.6
  3. 3
    Fa
    Fable 554.4

60 models

Vision Capabilities

Vision Capabilities Index

Average of the vision capability benchmarks published by the CAIS AI Dashboard.

  1. 1
    GPT-6 Astra logo
    GPT-6 Astra83.0
  2. 2
    Gemini 3.8 Flash logo
    Gemini 3.8 Flash69.9
  3. 3
    Fa
    Fable 5.168.9

48 models

Risk & Safety

Risk Index

Average of the risk and safety benchmarks published by the CAIS AI Dashboard; lower is better.

  1. 1
    Fa
    Fable 5.130.3
  2. 2
    Muse Spark 1.1 logo
    Muse Spark 1.132.5
  3. 3
    GPT-6 Astra logo
    GPT-6 Astra35.5

9 models

Automation

Automation

Remote Labor Index automation rates published by the CAIS AI Dashboard.

  1. 1
    Op
    Opus 5.521.3
  2. 2
    GPT-6 Astra logo
    GPT-6 Astra20.8
  3. 3
    Fa
    Fable 5.117.9

16 models

Reasoning

Intelligence Index

Composite Artificial Analysis index across mathematics, science, coding, and reasoning evaluations.

  1. 1
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)53.4
  2. 2
    Claude Fable 5.1 (xhigh with fallback) logo
    Claude Fable 5.1 (xhigh with fallback)53.2
  3. 3
    GPT-6 Astra (max) logo
    GPT-6 Astra (max)52.8

633 models

Code & Software Engineering

Coding Index

Artificial Analysis composite index for programming and software-engineering capability.

  1. 1
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)81.6
  2. 2
    Claude Fable 5.1 (xhigh with fallback) logo
    Claude Fable 5.1 (xhigh with fallback)80.7
  3. 3
    Claude Fable 5.1 (high with fallback) logo
    Claude Fable 5.1 (high with fallback)79.1

256 models

Task Planning & Knowledge Search

Agentic Index

Artificial Analysis composite index for tool use, planning, autonomy, and complex agentic workflows.

  1. 1
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)58.0
  2. 2
    Claude Fable 5.1 (xhigh with fallback) logo
    Claude Fable 5.1 (xhigh with fallback)57.1
  3. 3
    Claude Opus 5 (max) logo
    Claude Opus 5 (max)56.2

151 models

Knowledge

Omniscience Index

Artificial Analysis factual-knowledge and hallucination-resistance evaluation.

  1. 1
    GPT-6 Astra (high) logo
    GPT-6 Astra (high)43.7
  2. 2
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)43.5
  3. 3
    GPT-6 Astra (xhigh) logo
    GPT-6 Astra (xhigh)43.4

99 models

Knowledge

GPQA

Graduate-level, expert-written science questions evaluated by Artificial Analysis.

  1. 1
    GPT-6 Astra (xhigh) logo
    GPT-6 Astra (xhigh)96.3
  2. 2
    GPT-6 Astra (max) logo
    GPT-6 Astra (max)96.1
  3. 3
    Gemini 3.8 Flash (high) logo
    Gemini 3.8 Flash (high)95.3

612 models

Knowledge

Humanity's Last Exam

Broad expert-level reasoning and knowledge evaluation published by Artificial Analysis.

  1. 1
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)59.1
  2. 2
    Claude Fable 5.1 (xhigh with fallback) logo
    Claude Fable 5.1 (xhigh with fallback)58.7
  3. 3
    Claude Fable 5.1 (high with fallback) logo
    Claude Fable 5.1 (high with fallback)55.9

607 models

Instruction Following & Long Context

IFBench

Instruction-following benchmark results published by Artificial Analysis.

  1. 1
    Grok 4.3 (medium) logo
    Grok 4.3 (medium)83.3
  2. 2
    Grok 4.20 0309 logo
    Grok 4.20 030982.9
  3. 3
    MiniMax-M3 logo
    MiniMax-M382.9

450 models

Code & Software Engineering

SciCode

Scientific programming benchmark results published by Artificial Analysis.

  1. 1
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)63.1
  2. 2
    Claude Fable 5 (with fallback) logo
    Claude Fable 5 (with fallback)61.0
  3. 3
    Claude Fable 5.1 (xhigh with fallback) logo
    Claude Fable 5.1 (xhigh with fallback)60.9

167 models

Code & Software Engineering

Terminal-Bench 2.1

Agentic terminal and software-engineering results published by Artificial Analysis.

  1. 1
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)91.4
  2. 2
    Claude Fable 5.1 (xhigh with fallback) logo
    Claude Fable 5.1 (xhigh with fallback)91.0
  3. 3
    Claude Fable 5.1 (high with fallback) logo
    Claude Fable 5.1 (high with fallback)89.9

236 models

Code & Software Engineering

CritPt

Critical-points programming evaluation published by Artificial Analysis.

  1. 1
    GPT-5.6 Sol (max) logo
    GPT-5.6 Sol (max)32.3
  2. 2
    GPT-6 Astra (max) logo
    GPT-6 Astra (max)31.7
  3. 3
    GPT-6 Astra (xhigh) logo
    GPT-6 Astra (xhigh)31.4

520 models

Instruction Following & Long Context

Long Context Reasoning

Long-context reasoning results published by Artificial Analysis.

  1. 1
    Kimi K3 (max) logo
    Kimi K3 (max)88.7
  2. 2
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)85.3
  3. 3
    Claude Fable 5.1 (medium with fallback) logo
    Claude Fable 5.1 (medium with fallback)84.7

516 models

Task Planning & Knowledge Search

τ²-Bench

Tool-agent interaction benchmark results published by Artificial Analysis.

  1. 1
    JT
    JT-35B-Flash99.1
  2. 2
    GLM-5.2 (max) logo
    GLM-5.2 (max)99.1
  3. 3
    GLM-4.7-Flash logo
    GLM-4.7-Flash98.8

440 models

Scores are published snapshots and are only comparable within the same benchmark, version, metric, and evaluation configuration.