GenAiHub

LLMs

Browse 435 LLMs

LLMs Guide

Each model is ranked by its published Intelligence Index score and shown alongside its context window, measured output speed, blended price per million tokens, release date and whether its weights are open. Scores come from published benchmark results, not from our own testing.

76–100 of 435

LLM leaderboard

Rank reflects Intelligence Index

Click a column to sort
LLM leaderboard ranked before pagination using published benchmark results.
RankPosition based on the Intelligence Index. Tied scores share the same rank.ModelModel name and publisher matched to the published benchmark result.Composite score across reasoning, knowledge, mathematics, science, and coding evaluations.Composite score for programming and software-engineering capability.Composite score for tool use, planning, and autonomous task completion.Maximum number of tokens the model can process in one request.Median output generation throughput, measured in tokens per second.Blended USD cost per 1 million tokens using a 3:1 input-to-output ratio.Public release month and year for this model.Whether downloadable model weights are open or access is closed.
76
DeepSeek logo
DeepSeek V4 FlashDeepSeek
24.856.227.91M—$0.17Apr 2026Open source
77
So
Solar Open2 250BUpstage
24.745.0—1M——Aug 2026Open source
78
OpenAI logo
GPT-5.1OpenAI
24.749.4—272K121 tokens/s$3.44Nov 2025Closed source
79
OpenAI logo
GPT-5.4 miniOpenAI
24.656.119.7400K218 tokens/s$1.69Mar 2026Closed source
80
Xiaomi logo
MiMo-V2-OmniXiaomi
23.9——256K——Mar 2026Closed source
81
OpenAI logo
GPT-5.1 CodexOpenAI
23.7——400K—$3.44Nov 2025Closed source
82
Z AI logo
GLM 5V TurboZ AI
23.5——200K——Apr 2026Closed source
83
Kimi logo
Kimi K2.5Kimi
23.546.8—256K—$1.14Jan 2026Open source
84
Ne
Nemotron 3 UltraNVIDIA
23.449.321.7262.1K168 tokens/s$1.1Jun 2026Open source
85
MiniMax logo
MiniMax-M2.7MiniMax
23.252.616.8204.8K73 tokens/s$0.52Mar 2026Open source
86
OpenAI logo
GPT-5OpenAI
23.037.8—400K99 tokens/s$3.44Aug 2025Closed source
87
Alibaba logo
Qwen3.5 27BAlibaba
22.9——262.1K77 tokens/s$0.83Feb 2026Open source
88
Anthropic logo
Claude 4.1 OpusAnthropic
22.8——200K—$30Aug 2025Closed source
89
MiniMax logo
MiniMax-M2.5MiniMax
22.8——204.8K96 tokens/s$0.52Feb 2026Open source
90
Hy
Hy3-previewTencent
22.7——256K—$0.1Apr 2026Open source
91
A.
A.X-K2SK Telecom
22.738.8—262.1K——Aug 2026Open source
92
Google logo
Gemini 3.5 Flash-LiteGoogle
22.749.315.91M363 tokens/s$0.85Jul 2026Closed source
93
SpaceXAI logo
Grok 4SpaceXAI
22.5——256K—$6Jul 2025Closed source
94
Xiaomi logo
MiMo-V2-FlashXiaomi
22.449.8—256K——Dec 2025Open source
95
Xiaomi logo
MiMo-V2.5Xiaomi
22.356.817.41M65 tokens/s$0.18Apr 2026Open source
96
Alibaba logo
Qwen3.6 35B A3BAlibaba
22.341.915.0262.1K120 tokens/s$0.84Apr 2026Open source
97
Z AI logo
GLM-4.7Z AI
22.245.3—200K97 tokens/s$1Dec 2025Open source
98
Kimi logo
Kimi K2 ThinkingKimi
22.0——256K121 tokens/s$1.08Nov 2025Open source
99
Alibaba logo
Qwen3.6 27BAlibaba
21.953.720.1262.1K56 tokens/s$1.35Apr 2026Open source
100
OpenAI logo
o3-proOpenAI
21.9——200K—$35Jun 2025Closed source

Rank uses Intelligence Index. Sort any index, context, speed, pricing, release date, or license; missing data always appears last.

All benchmark leaderboards

Explore every published snapshot across coding, agentic search, reasoning, instruction following, long context, and safety evaluations.

Browse by category

31 of 31 benchmarks

Code & Software Engineering

SWE-bench Pro

A long-horizon repository benchmark built to test realistic software engineering work. Its scores need a task-quality and setup check before they support a coding-agent decision.

  1. 1
    Claude Opus 5.5 logo
    Claude Opus 5.589.9
  2. 2
    Claude Sonnet 5.5 logo
    Claude Sonnet 5.581.3
  3. 3
    Claude Fable 5.1 logo
    Claude Fable 5.181.2

78 models

Code & Software Engineering

SWE Multilingual

A multilingual software-engineering benchmark for real-world code issue resolution across multiple programming languages.

  1. 1
    Claude Opus 5.5 logo
    Claude Opus 5.593.9
  2. 2
    Claude Sonnet 5.5 logo
    Claude Sonnet 5.590.3
  3. 3
    Claude Opus 5 logo
    Claude Opus 589.5

48 models

Financial & Business Operations

AA Tau3 Banking

An independently evaluated Tau3 banking benchmark from Artificial Analysis.

  1. 1
    Grok 4.6 logo
    Grok 4.650.7
  2. 2
    Muse Spark 1.3 logo
    Muse Spark 1.350.5
  3. 3
    GLM-5.3 logo
    GLM-5.350.3

14 models

Task Planning & Knowledge Search

SkillsBench

How important are skills for agents?

  1. 1
    DeepSeek V4.1 Flash logo
    DeepSeek V4.1 Flash69.8
  2. 2
    Grok 4.5 logo
    Grok 4.566.0
  3. 3
    Gemini 3.7 Flash logo
    Gemini 3.7 Flash65.9

35 models

Task Planning & Knowledge Search

WideResearch

A broad research-agent benchmark for open-ended information gathering, synthesis, and answer construction across wide search spaces.

  1. 1
    Me
    Mercury 292.3
  2. 2
    Hy
    Hy4 preview83.9
  3. 3
    Qwen3.8 Max logo
    Qwen3.8 Max81.9

16 models

Task Planning & Knowledge Search

BrowseComp

A benchmark for web-browsing agents that must search, inspect sources, gather evidence, and return the correct answer to research-oriented questions.

  1. 1
    At
    Atria Dawn Preview92.5
  2. 2
    GPT-5.6 Sol logo
    GPT-5.6 Sol92.2
  3. 3
    GPT-6 Astra logo
    GPT-6 Astra91.5

47 models

Instruction Following & Long Context

MRCR v2 128K-256K

MRCR v2 slice focused on very long contexts at 128K-256K lengths.

  1. 1
    GPT-5.5 logo
    GPT-5.587.5
  2. 2
    Claude Opus 4.7 (Adaptive) logo
    Claude Opus 4.7 (Adaptive)59.2

2 models

Code & Software Engineering

SWE-bench Verified

Human-validated software engineering issues from real-world Python repositories.

  1. 1
    Claude Opus 4.7 (max) logo
    Claude Opus 4.7 (max)83.5
  2. 2
    GPT-5.5 (xhigh) logo
    GPT-5.5 (xhigh)80.6
  3. 3
    Gemini 3.5 Flash (high) logo
    Gemini 3.5 Flash (high)79.3

33 models

Knowledge

SimpleQA Verified

Factoid questions spanning politics, science, technology, art, sports, geography, and music.

  1. 1
    GPT-6 Astra (max) logo
    GPT-6 Astra (max)75.6
  2. 2
    GPT-6.1 Sol (max) logo
    GPT-6.1 Sol (max)73.9
  3. 3
    Gemini 3.1 Pro Preview (high) logo
    Gemini 3.1 Pro Preview (high)73.5

86 models

Reasoning

ARC-AGI-2

Abstract reasoning and pattern generalization on grid-based tasks.

  1. 1
    GPT-6 Astra (max) logo
    GPT-6 Astra (max)95.0
  2. 2
    GPT-6.1 Sol (max) logo
    GPT-6.1 Sol (max)94.2
  3. 3
    GPT-6 Astra (xhigh) logo
    GPT-6 Astra (xhigh)93.3

235 models

Math and science

FrontierMath Tier 4 (v2)

Exceptionally difficult research-level mathematics problems.

  1. 1
    GPT-6.1 Sol (max) logo
    GPT-6.1 Sol (max)100.0
  2. 2
    GPT-6 Astra (high) logo
    GPT-6 Astra (high)97.6
  3. 3
    GPT-6 Astra (xhigh) logo
    GPT-6 Astra (xhigh)97.6

70 models

Task Planning & Knowledge Search

MCP Atlas

Real-world multi-step tool-use evaluation through the Model Context Protocol.

  1. 1
    Muse Spark 1.1 logo
    Muse Spark 1.188.1
  2. 2
    Fa
    Fable 5.187.2
  3. 3
    claude-opus-5 (xhigh) logo
    claude-opus-5 (xhigh)85.8

34 models

Cybersecurity & Risk

CyberGym

Real-world cybersecurity evaluation of AI agents reproducing vulnerabilities with working proof-of-concept tests.

  1. 1
    Multi-model (DeepSeek-V4.1-Flash, Abliterated GLM-5.2) logo
    Multi-model (DeepSeek-V4.1-Flash, Abliterated GLM-5.2)99.2
  2. 2
    Multi-model (Creation Model, DeepSeek-V4-Pro, Qwen 3.8 Max) logo
    Multi-model (Creation Model, DeepSeek-V4-Pro, Qwen 3.8 Max)98.5
  3. 3
    Multi-model (DeepSeek-V4-Flash, OrcaCyber-Zero-1.0) logo
    Multi-model (DeepSeek-V4-Flash, OrcaCyber-Zero-1.0)98.0

77 models

Code & Software Engineering

Terminal-Bench 3.0

The official leaderboard for Terminal-Bench 3.0.

  1. 1
    Op
    Opus 542.7
  2. 2
    GPT-5.6 Sol logo
    GPT-5.6 Sol34.6
  3. 3
    Fa
    Fable 534.0

12 models

Instruction Following & Long Context

Humanity's Last Exam

Expert-level questions across mathematics, science, and humanities, as published by the CAIS AI Dashboard.

  1. 1
    Fa
    Fable 5.154.6
  2. 2
    GPT-6 Astra logo
    GPT-6 Astra53.6
  3. 3
    Fa
    Fable 552.7

60 models

Text Capabilities

Text Capabilities Index

Average of the text capability benchmarks published by the CAIS AI Dashboard.

  1. 1
    GPT-6 Astra logo
    GPT-6 Astra63.9
  2. 2
    Fa
    Fable 5.155.6
  3. 3
    Fa
    Fable 554.4

60 models

Vision Capabilities

Vision Capabilities Index

Average of the vision capability benchmarks published by the CAIS AI Dashboard.

  1. 1
    GPT-6 Astra logo
    GPT-6 Astra83.0
  2. 2
    Gemini 3.8 Flash logo
    Gemini 3.8 Flash69.9
  3. 3
    Fa
    Fable 5.168.9

48 models

Risk & Safety

Risk Index

Average of the risk and safety benchmarks published by the CAIS AI Dashboard; lower is better.

  1. 1
    Fa
    Fable 5.130.3
  2. 2
    Muse Spark 1.1 logo
    Muse Spark 1.132.5
  3. 3
    GPT-6 Astra logo
    GPT-6 Astra35.5

9 models

Automation

Automation

Remote Labor Index automation rates published by the CAIS AI Dashboard.

  1. 1
    Op
    Opus 5.521.3
  2. 2
    GPT-6 Astra logo
    GPT-6 Astra20.8
  3. 3
    Fa
    Fable 5.117.9

16 models

Reasoning

Intelligence Index

Composite Artificial Analysis index across mathematics, science, coding, and reasoning evaluations.

  1. 1
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)53.4
  2. 2
    Claude Fable 5.1 (xhigh with fallback) logo
    Claude Fable 5.1 (xhigh with fallback)53.2
  3. 3
    GPT-6 Astra (max) logo
    GPT-6 Astra (max)52.8

633 models

Code & Software Engineering

Coding Index

Artificial Analysis composite index for programming and software-engineering capability.

  1. 1
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)81.6
  2. 2
    Claude Fable 5.1 (xhigh with fallback) logo
    Claude Fable 5.1 (xhigh with fallback)80.7
  3. 3
    Claude Fable 5.1 (high with fallback) logo
    Claude Fable 5.1 (high with fallback)79.1

256 models

Task Planning & Knowledge Search

Agentic Index

Artificial Analysis composite index for tool use, planning, autonomy, and complex agentic workflows.

  1. 1
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)58.0
  2. 2
    Claude Fable 5.1 (xhigh with fallback) logo
    Claude Fable 5.1 (xhigh with fallback)57.1
  3. 3
    Claude Opus 5 (max) logo
    Claude Opus 5 (max)56.2

151 models

Knowledge

Omniscience Index

Artificial Analysis factual-knowledge and hallucination-resistance evaluation.

  1. 1
    GPT-6 Astra (high) logo
    GPT-6 Astra (high)43.7
  2. 2
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)43.5
  3. 3
    GPT-6 Astra (xhigh) logo
    GPT-6 Astra (xhigh)43.4

99 models

Knowledge

GPQA

Graduate-level, expert-written science questions evaluated by Artificial Analysis.

  1. 1
    GPT-6 Astra (xhigh) logo
    GPT-6 Astra (xhigh)96.3
  2. 2
    GPT-6 Astra (max) logo
    GPT-6 Astra (max)96.1
  3. 3
    Gemini 3.8 Flash (high) logo
    Gemini 3.8 Flash (high)95.3

612 models

Knowledge

Humanity's Last Exam

Broad expert-level reasoning and knowledge evaluation published by Artificial Analysis.

  1. 1
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)59.1
  2. 2
    Claude Fable 5.1 (xhigh with fallback) logo
    Claude Fable 5.1 (xhigh with fallback)58.7
  3. 3
    Claude Fable 5.1 (high with fallback) logo
    Claude Fable 5.1 (high with fallback)55.9

607 models

Instruction Following & Long Context

IFBench

Instruction-following benchmark results published by Artificial Analysis.

  1. 1
    Grok 4.3 (medium) logo
    Grok 4.3 (medium)83.3
  2. 2
    Grok 4.20 0309 logo
    Grok 4.20 030982.9
  3. 3
    MiniMax-M3 logo
    MiniMax-M382.9

450 models

Code & Software Engineering

SciCode

Scientific programming benchmark results published by Artificial Analysis.

  1. 1
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)63.1
  2. 2
    Claude Fable 5 (with fallback) logo
    Claude Fable 5 (with fallback)61.0
  3. 3
    Claude Fable 5.1 (xhigh with fallback) logo
    Claude Fable 5.1 (xhigh with fallback)60.9

167 models

Code & Software Engineering

Terminal-Bench 2.1

Agentic terminal and software-engineering results published by Artificial Analysis.

  1. 1
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)91.4
  2. 2
    Claude Fable 5.1 (xhigh with fallback) logo
    Claude Fable 5.1 (xhigh with fallback)91.0
  3. 3
    Claude Fable 5.1 (high with fallback) logo
    Claude Fable 5.1 (high with fallback)89.9

236 models

Code & Software Engineering

CritPt

Critical-points programming evaluation published by Artificial Analysis.

  1. 1
    GPT-5.6 Sol (max) logo
    GPT-5.6 Sol (max)32.3
  2. 2
    GPT-6 Astra (max) logo
    GPT-6 Astra (max)31.7
  3. 3
    GPT-6 Astra (xhigh) logo
    GPT-6 Astra (xhigh)31.4

520 models

Instruction Following & Long Context

Long Context Reasoning

Long-context reasoning results published by Artificial Analysis.

  1. 1
    Kimi K3 (max) logo
    Kimi K3 (max)88.7
  2. 2
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)85.3
  3. 3
    Claude Fable 5.1 (medium with fallback) logo
    Claude Fable 5.1 (medium with fallback)84.7

516 models

Task Planning & Knowledge Search

τ²-Bench

Tool-agent interaction benchmark results published by Artificial Analysis.

  1. 1
    JT
    JT-35B-Flash99.1
  2. 2
    GLM-5.2 (max) logo
    GLM-5.2 (max)99.1
  3. 3
    GLM-4.7-Flash logo
    GLM-4.7-Flash98.8

440 models

Scores are published snapshots and are only comparable within the same benchmark, version, metric, and evaluation configuration.