GenAiHub

LLMs

Browse 435 LLMs

LLMs Guide

Each model is ranked by its published Intelligence Index score and shown alongside its context window, measured output speed, blended price per million tokens, release date and whether its weights are open. Scores come from published benchmark results, not from our own testing.

1–25 of 435

LLM leaderboard

Rank reflects Intelligence Index

Click a column to sort
LLM leaderboard ranked before pagination using published benchmark results.
RankPosition based on the Intelligence Index. Tied scores share the same rank.ModelModel name and publisher matched to the published benchmark result.Composite score across reasoning, knowledge, mathematics, science, and coding evaluations.Composite score for programming and software-engineering capability.Composite score for tool use, planning, and autonomous task completion.Maximum number of tokens the model can process in one request.Median output generation throughput, measured in tokens per second.Blended USD cost per 1 million tokens using a 3:1 input-to-output ratio.Public release month and year for this model.Whether downloadable model weights are open or access is closed.
1
Anthropic logo
Claude Fable 5.1Anthropic
53.481.658.01M69 tokens/s$20Sep 2026Closed source
2
OpenAI logo
GPT-6 AstraOpenAI
52.877.151.51M56 tokens/s$20Sep 2026Closed source
3
Anthropic logo
Claude Opus 5Anthropic
50.778.056.21M54 tokens/s$10Jul 2026Closed source
4
Anthropic logo
Claude Fable 5Anthropic
49.776.551.01M64 tokens/s$20Jun 2026Closed source
5
Meta logo
Muse Spark 1.3Meta
48.276.555.71M232 tokens/s$2Sep 2026Closed source
6
OpenAI logo
GPT-5.6 SolOpenAI
47.178.350.51M68 tokens/s$8Jul 2026Closed source
7
Z AI logo
GLM-5.3Z AI
44.974.853.41M70 tokens/s$2.15Aug 2026Open source
8
SpaceXAI logo
Grok 4.6SpaceXAI
44.476.853.4500K58 tokens/s$3Aug 2026Closed source
9
Kimi logo
Kimi K3Kimi
43.876.250.61M41 tokens/s$6Jul 2026Open source
10
OpenAI logo
GPT-5.6 TerraOpenAI
42.376.743.71M116 tokens/s$4.5Jul 2026Closed source
11
Alibaba logo
Qwen3.8-Flash-NextAlibaba
42.273.1—256K72 tokens/s$0.23Aug 2026Open source
12
Anthropic logo
Claude Opus 4.8Anthropic
42.074.342.61M59 tokens/s$10May 2026Closed source
13
Z AI logo
GLM-5.3-FlashZ AI
41.971.551.21M71 tokens/s$0.24Aug 2026Open source
14
Google logo
Gemini 3.8 FlashGoogle
41.276.341.11M281 tokens/s$1.5Sep 2026Closed source
15
Anthropic logo
Claude Opus 4.7Anthropic
40.773.639.51M50 tokens/s$10Apr 2026Closed source
16
Alibaba logo
Qwen3.8 MaxAlibaba
40.371.849.61M40 tokens/s$3Aug 2026Closed source
17
Alibaba logo
Qwen3.8 2.4T A95BAlibaba
40.071.950.4983.6K41 tokens/s$3Aug 2026Open source
18
Meta logo
Muse Spark 1.2Meta
39.872.244.01M238 tokens/s$2Aug 2026Closed source
19
Google logo
Gemini 3.7 FlashGoogle
39.676.136.41M311 tokens/s$1.5Aug 2026Closed source
20
SpaceXAI logo
Grok 4.5SpaceXAI
39.172.442.1500K59 tokens/s$3Jul 2026Closed source
21
OpenAI logo
GPT-5.4OpenAI
39.071.1—1.1M158 tokens/s$5.63Mar 2026Closed source
22
Z AI logo
GLM-5.2Z AI
38.668.839.41M63 tokens/s$2.15Jun 2026Open source
23
OpenAI logo
GPT-5.5OpenAI
38.674.937.3922K90 tokens/s$11.25Apr 2026Closed source
24
Anthropic logo
Claude Sonnet 5Anthropic
38.471.544.31M79 tokens/s$4Jun 2026Closed source
25
OpenAI logo
GPT-5.6 LunaOpenAI
37.571.442.71M120 tokens/s$0.45Jul 2026Closed source

Rank uses Intelligence Index. Sort any index, context, speed, pricing, release date, or license; missing data always appears last.

All benchmark leaderboards

Explore every published snapshot across coding, agentic search, reasoning, instruction following, long context, and safety evaluations.

Browse by category

31 of 31 benchmarks

Code & Software Engineering

SWE-bench Pro

A long-horizon repository benchmark built to test realistic software engineering work. Its scores need a task-quality and setup check before they support a coding-agent decision.

  1. 1
    Claude Opus 5.5 logo
    Claude Opus 5.589.9
  2. 2
    Claude Sonnet 5.5 logo
    Claude Sonnet 5.581.3
  3. 3
    Claude Fable 5.1 logo
    Claude Fable 5.181.2

78 models

Code & Software Engineering

SWE Multilingual

A multilingual software-engineering benchmark for real-world code issue resolution across multiple programming languages.

  1. 1
    Claude Opus 5.5 logo
    Claude Opus 5.593.9
  2. 2
    Claude Sonnet 5.5 logo
    Claude Sonnet 5.590.3
  3. 3
    Claude Opus 5 logo
    Claude Opus 589.5

48 models

Financial & Business Operations

AA Tau3 Banking

An independently evaluated Tau3 banking benchmark from Artificial Analysis.

  1. 1
    Grok 4.6 logo
    Grok 4.650.7
  2. 2
    Muse Spark 1.3 logo
    Muse Spark 1.350.5
  3. 3
    GLM-5.3 logo
    GLM-5.350.3

14 models

Task Planning & Knowledge Search

SkillsBench

How important are skills for agents?

  1. 1
    DeepSeek V4.1 Flash logo
    DeepSeek V4.1 Flash69.8
  2. 2
    Grok 4.5 logo
    Grok 4.566.0
  3. 3
    Gemini 3.7 Flash logo
    Gemini 3.7 Flash65.9

35 models

Task Planning & Knowledge Search

WideResearch

A broad research-agent benchmark for open-ended information gathering, synthesis, and answer construction across wide search spaces.

  1. 1
    Me
    Mercury 292.3
  2. 2
    Hy
    Hy4 preview83.9
  3. 3
    Qwen3.8 Max logo
    Qwen3.8 Max81.9

16 models

Task Planning & Knowledge Search

BrowseComp

A benchmark for web-browsing agents that must search, inspect sources, gather evidence, and return the correct answer to research-oriented questions.

  1. 1
    At
    Atria Dawn Preview92.5
  2. 2
    GPT-5.6 Sol logo
    GPT-5.6 Sol92.2
  3. 3
    GPT-6 Astra logo
    GPT-6 Astra91.5

47 models

Instruction Following & Long Context

MRCR v2 128K-256K

MRCR v2 slice focused on very long contexts at 128K-256K lengths.

  1. 1
    GPT-5.5 logo
    GPT-5.587.5
  2. 2
    Claude Opus 4.7 (Adaptive) logo
    Claude Opus 4.7 (Adaptive)59.2

2 models

Code & Software Engineering

SWE-bench Verified

Human-validated software engineering issues from real-world Python repositories.

  1. 1
    Claude Opus 4.7 (max) logo
    Claude Opus 4.7 (max)83.5
  2. 2
    GPT-5.5 (xhigh) logo
    GPT-5.5 (xhigh)80.6
  3. 3
    Gemini 3.5 Flash (high) logo
    Gemini 3.5 Flash (high)79.3

33 models

Knowledge

SimpleQA Verified

Factoid questions spanning politics, science, technology, art, sports, geography, and music.

  1. 1
    GPT-6 Astra (max) logo
    GPT-6 Astra (max)75.6
  2. 2
    GPT-6.1 Sol (max) logo
    GPT-6.1 Sol (max)73.9
  3. 3
    Gemini 3.1 Pro Preview (high) logo
    Gemini 3.1 Pro Preview (high)73.5

86 models

Reasoning

ARC-AGI-2

Abstract reasoning and pattern generalization on grid-based tasks.

  1. 1
    GPT-6 Astra (max) logo
    GPT-6 Astra (max)95.0
  2. 2
    GPT-6.1 Sol (max) logo
    GPT-6.1 Sol (max)94.2
  3. 3
    GPT-6 Astra (xhigh) logo
    GPT-6 Astra (xhigh)93.3

235 models

Math and science

FrontierMath Tier 4 (v2)

Exceptionally difficult research-level mathematics problems.

  1. 1
    GPT-6.1 Sol (max) logo
    GPT-6.1 Sol (max)100.0
  2. 2
    GPT-6 Astra (high) logo
    GPT-6 Astra (high)97.6
  3. 3
    GPT-6 Astra (xhigh) logo
    GPT-6 Astra (xhigh)97.6

70 models

Task Planning & Knowledge Search

MCP Atlas

Real-world multi-step tool-use evaluation through the Model Context Protocol.

  1. 1
    Muse Spark 1.1 logo
    Muse Spark 1.188.1
  2. 2
    Fa
    Fable 5.187.2
  3. 3
    claude-opus-5 (xhigh) logo
    claude-opus-5 (xhigh)85.8

34 models

Cybersecurity & Risk

CyberGym

Real-world cybersecurity evaluation of AI agents reproducing vulnerabilities with working proof-of-concept tests.

  1. 1
    Multi-model (DeepSeek-V4.1-Flash, Abliterated GLM-5.2) logo
    Multi-model (DeepSeek-V4.1-Flash, Abliterated GLM-5.2)99.2
  2. 2
    Multi-model (Creation Model, DeepSeek-V4-Pro, Qwen 3.8 Max) logo
    Multi-model (Creation Model, DeepSeek-V4-Pro, Qwen 3.8 Max)98.5
  3. 3
    Multi-model (DeepSeek-V4-Flash, OrcaCyber-Zero-1.0) logo
    Multi-model (DeepSeek-V4-Flash, OrcaCyber-Zero-1.0)98.0

77 models

Code & Software Engineering

Terminal-Bench 3.0

The official leaderboard for Terminal-Bench 3.0.

  1. 1
    Op
    Opus 542.7
  2. 2
    GPT-5.6 Sol logo
    GPT-5.6 Sol34.6
  3. 3
    Fa
    Fable 534.0

12 models

Instruction Following & Long Context

Humanity's Last Exam

Expert-level questions across mathematics, science, and humanities, as published by the CAIS AI Dashboard.

  1. 1
    Fa
    Fable 5.154.6
  2. 2
    GPT-6 Astra logo
    GPT-6 Astra53.6
  3. 3
    Fa
    Fable 552.7

60 models

Text Capabilities

Text Capabilities Index

Average of the text capability benchmarks published by the CAIS AI Dashboard.

  1. 1
    GPT-6 Astra logo
    GPT-6 Astra63.9
  2. 2
    Fa
    Fable 5.155.6
  3. 3
    Fa
    Fable 554.4

60 models

Vision Capabilities

Vision Capabilities Index

Average of the vision capability benchmarks published by the CAIS AI Dashboard.

  1. 1
    GPT-6 Astra logo
    GPT-6 Astra83.0
  2. 2
    Gemini 3.8 Flash logo
    Gemini 3.8 Flash69.9
  3. 3
    Fa
    Fable 5.168.9

48 models

Risk & Safety

Risk Index

Average of the risk and safety benchmarks published by the CAIS AI Dashboard; lower is better.

  1. 1
    Fa
    Fable 5.130.3
  2. 2
    Muse Spark 1.1 logo
    Muse Spark 1.132.5
  3. 3
    GPT-6 Astra logo
    GPT-6 Astra35.5

9 models

Automation

Automation

Remote Labor Index automation rates published by the CAIS AI Dashboard.

  1. 1
    Op
    Opus 5.521.3
  2. 2
    GPT-6 Astra logo
    GPT-6 Astra20.8
  3. 3
    Fa
    Fable 5.117.9

16 models

Reasoning

Intelligence Index

Composite Artificial Analysis index across mathematics, science, coding, and reasoning evaluations.

  1. 1
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)53.4
  2. 2
    Claude Fable 5.1 (xhigh with fallback) logo
    Claude Fable 5.1 (xhigh with fallback)53.2
  3. 3
    GPT-6 Astra (max) logo
    GPT-6 Astra (max)52.8

633 models

Code & Software Engineering

Coding Index

Artificial Analysis composite index for programming and software-engineering capability.

  1. 1
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)81.6
  2. 2
    Claude Fable 5.1 (xhigh with fallback) logo
    Claude Fable 5.1 (xhigh with fallback)80.7
  3. 3
    Claude Fable 5.1 (high with fallback) logo
    Claude Fable 5.1 (high with fallback)79.1

256 models

Task Planning & Knowledge Search

Agentic Index

Artificial Analysis composite index for tool use, planning, autonomy, and complex agentic workflows.

  1. 1
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)58.0
  2. 2
    Claude Fable 5.1 (xhigh with fallback) logo
    Claude Fable 5.1 (xhigh with fallback)57.1
  3. 3
    Claude Opus 5 (max) logo
    Claude Opus 5 (max)56.2

151 models

Knowledge

Omniscience Index

Artificial Analysis factual-knowledge and hallucination-resistance evaluation.

  1. 1
    GPT-6 Astra (high) logo
    GPT-6 Astra (high)43.7
  2. 2
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)43.5
  3. 3
    GPT-6 Astra (xhigh) logo
    GPT-6 Astra (xhigh)43.4

99 models

Knowledge

GPQA

Graduate-level, expert-written science questions evaluated by Artificial Analysis.

  1. 1
    GPT-6 Astra (xhigh) logo
    GPT-6 Astra (xhigh)96.3
  2. 2
    GPT-6 Astra (max) logo
    GPT-6 Astra (max)96.1
  3. 3
    Gemini 3.8 Flash (high) logo
    Gemini 3.8 Flash (high)95.3

612 models

Knowledge

Humanity's Last Exam

Broad expert-level reasoning and knowledge evaluation published by Artificial Analysis.

  1. 1
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)59.1
  2. 2
    Claude Fable 5.1 (xhigh with fallback) logo
    Claude Fable 5.1 (xhigh with fallback)58.7
  3. 3
    Claude Fable 5.1 (high with fallback) logo
    Claude Fable 5.1 (high with fallback)55.9

607 models

Instruction Following & Long Context

IFBench

Instruction-following benchmark results published by Artificial Analysis.

  1. 1
    Grok 4.3 (medium) logo
    Grok 4.3 (medium)83.3
  2. 2
    Grok 4.20 0309 logo
    Grok 4.20 030982.9
  3. 3
    MiniMax-M3 logo
    MiniMax-M382.9

450 models

Code & Software Engineering

SciCode

Scientific programming benchmark results published by Artificial Analysis.

  1. 1
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)63.1
  2. 2
    Claude Fable 5 (with fallback) logo
    Claude Fable 5 (with fallback)61.0
  3. 3
    Claude Fable 5.1 (xhigh with fallback) logo
    Claude Fable 5.1 (xhigh with fallback)60.9

167 models

Code & Software Engineering

Terminal-Bench 2.1

Agentic terminal and software-engineering results published by Artificial Analysis.

  1. 1
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)91.4
  2. 2
    Claude Fable 5.1 (xhigh with fallback) logo
    Claude Fable 5.1 (xhigh with fallback)91.0
  3. 3
    Claude Fable 5.1 (high with fallback) logo
    Claude Fable 5.1 (high with fallback)89.9

236 models

Code & Software Engineering

CritPt

Critical-points programming evaluation published by Artificial Analysis.

  1. 1
    GPT-5.6 Sol (max) logo
    GPT-5.6 Sol (max)32.3
  2. 2
    GPT-6 Astra (max) logo
    GPT-6 Astra (max)31.7
  3. 3
    GPT-6 Astra (xhigh) logo
    GPT-6 Astra (xhigh)31.4

520 models

Instruction Following & Long Context

Long Context Reasoning

Long-context reasoning results published by Artificial Analysis.

  1. 1
    Kimi K3 (max) logo
    Kimi K3 (max)88.7
  2. 2
    Claude Fable 5.1 (max with fallback) logo
    Claude Fable 5.1 (max with fallback)85.3
  3. 3
    Claude Fable 5.1 (medium with fallback) logo
    Claude Fable 5.1 (medium with fallback)84.7

516 models

Task Planning & Knowledge Search

τ²-Bench

Tool-agent interaction benchmark results published by Artificial Analysis.

  1. 1
    JT
    JT-35B-Flash99.1
  2. 2
    GLM-5.2 (max) logo
    GLM-5.2 (max)99.1
  3. 3
    GLM-4.7-Flash logo
    GLM-4.7-Flash98.8

440 models

Scores are published snapshots and are only comparable within the same benchmark, version, metric, and evaluation configuration.