GenAiHub
Knowledge
Epoch live leaderboard

SimpleQA Verified

Factoid questions spanning politics, science, technology, art, sports, geography, and music.

Metric

Accuracy

Results shown

86

Publisher

Epoch AI live leaderboard

Snapshot

Fetched 2026-10-08T23:19:32.614Z

86 of 86 models

86 of 86 models

Higher is better

GPT-6 Astra (max) logo
GPT-6 Astra (max)
GPT-6.1 Sol (max) logo
GPT-6.1 Sol (max)
Gemini 3.1 Pro Preview (high) logo
Gemini 3.1 Pro Preview (high)
Claude Opus 5.5 (max) logo
Claude Opus 5.5 (max)
Claude Fable 5.1 (max) logo
Claude Fable 5.1 (max)
Claude Fable 5 (xhigh) logo
Claude Fable 5 (xhigh)
Gemini 3.8 Flash (high) logo
Gemini 3.8 Flash (high)
GPT-5.6 Sol (max) logo
GPT-5.6 Sol (max)
Gemini 3.7 Flash (high) logo
Gemini 3.7 Flash (high)
Gemini 3 Flash Preview (high) logo
Gemini 3 Flash Preview (high)
Gemini 3.5 Flash (high) logo
Gemini 3.5 Flash (high)
Gemini 3.6 Flash (high) logo
Gemini 3.6 Flash (high)
GPT-5.5 (xhigh) logo
GPT-5.5 (xhigh)
GPT-6 Sol (max) logo
GPT-6 Sol (max)
Muse Spark 1.2 (xhigh) logo
Muse Spark 1.2 (xhigh)
Claude Opus 5 (max) logo
Claude Opus 5 (max)
Muse Spark 1.1 logo
Muse Spark 1.1
Grok 4.7 (xhigh) logo
Grok 4.7 (xhigh)
Qwen3.7 Max logo
Qwen3.7 Max
Claude Opus 4.8 (max) logo
Claude Opus 4.8 (max)
DeepSeek V4 Pro 0813 (max) logo
DeepSeek V4 Pro 0813 (max)
Qwen3.6 Max Preview (thinking) logo
Qwen3.6 Max Preview (thinking)
Claude Opus 4.7 (xhigh) logo
Claude Opus 4.7 (xhigh)
Kimi K3 (Max) logo
Kimi K3 (Max)
GPT-5 (high) logo
GPT-5 (high)
o3 (high) logo
o3 (high)
Grok 4.6 (high) logo
Grok 4.6 (high)
Grok 4.6 (xhigh) logo
Grok 4.6 (xhigh)
Qwen3 Max logo
Qwen3 Max
Grok 4.5 (high) logo
Grok 4.5 (high)
GPT-5.1 (high) logo
GPT-5.1 (high)
Qwen3.8 Max (0902) (xhigh) logo
Qwen3.8 Max (0902) (xhigh)
Claude Opus 4.6 (max) logo
Claude Opus 4.6 (max)
DeepSeek v4 (max) logo
DeepSeek v4 (max)
Claude Sonnet 5.5 (max) logo
Claude Sonnet 5.5 (max)
GPT-5.4 Pro (xhigh) logo
GPT-5.4 Pro (xhigh)
Qwen3.8 Max (xhigh) logo
Qwen3.8 Max (xhigh)
Claude Opus 4.5 (32k thinking) logo
Claude Opus 4.5 (32k thinking)
GPT-5.4 (xhigh) logo
GPT-5.4 (xhigh)
Qwen3.6 Plus (thinking) logo
Qwen3.6 Plus (thinking)
GPT-5.6 Terra (max) logo
GPT-5.6 Terra (max)
GPT-6 Luna (max) logo
GPT-6 Luna (max)
o1 (high) logo
o1 (high)
GLM-5.3 (max) logo
GLM-5.3 (max)
GPT-5.6 Luna (max) logo
GPT-5.6 Luna (max)
Qwen3-235B-A22B (Jul 2025) logo
Qwen3-235B-A22B (Jul 2025)
In
Inkling (xhigh)
GPT-5.2 (xhigh) logo
GPT-5.2 (xhigh)
Kimi K2.7 Code logo
Kimi K2.7 Code
Claude Sonnet 4.6 (high) logo
Claude Sonnet 4.6 (high)
Kimi K2.6 logo
Kimi K2.6
Kimi K2.5 logo
Kimi K2.5
GPT-5.2 (high) logo
GPT-5.2 (high)
GLM-5.2 (max) logo
GLM-5.2 (max)
GLM-5.1 logo
GLM-5.1
Claude Sonnet 5 (max) logo
Claude Sonnet 5 (max)
DeepSeek V4 Flash 0731 (max) logo
DeepSeek V4 Flash 0731 (max)
Grok 4.3 Beta logo
Grok 4.3 Beta
Claude Sonnet 5 (xhigh) logo
Claude Sonnet 5 (xhigh)
GPT-5.2 (low) logo
GPT-5.2 (low)
Claude Sonnet 4.6 (max) logo
Claude Sonnet 4.6 (max)
GPT-5.2 (medium) logo
GPT-5.2 (medium)
GLM-4.7 logo
GLM-4.7
GPT-4.1 logo
GPT-4.1
Claude Sonnet 4.5 (59k thinking) logo
Claude Sonnet 4.5 (59k thinking)
Grok 4.20 logo
Grok 4.20
GPT-5.4 mini (high) logo
GPT-5.4 mini (high)
GPT-4o (Aug 2024) logo
GPT-4o (Aug 2024)
Qwen3.5 Plus (thinking) logo
Qwen3.5 Plus (thinking)
Claude Sonnet 4.5 (no thinking) logo
Claude Sonnet 4.5 (no thinking)
GPT-5 mini (high) logo
GPT-5 mini (high)
Qwen3.5 Flash (thinking) logo
Qwen3.5 Flash (thinking)
o4-mini (low) logo
o4-mini (low)
In
Inkling Small (xhigh)
o4-mini (high) logo
o4-mini (high)
Qwen3.6 Flash (thinking) logo
Qwen3.6 Flash (thinking)
o3-mini (high) logo
o3-mini (high)
Claude Haiku 4.5 logo
Claude Haiku 4.5
GPT-4.1 mini logo
GPT-4.1 mini
Claude 3 Opus logo
Claude 3 Opus
Claude Haiku 4.5 (32k thinking) logo
Claude Haiku 4.5 (32k thinking)
GPT-5.4 nano (high) logo
GPT-5.4 nano (high)
GPT-5 nano (high) logo
GPT-5 nano (high)
Ge
Gemma 4 31B IT
GPT-4o mini logo
GPT-4o mini
GPT-4.1 nano logo
GPT-4.1 nano
List bars visualize relative performance. Plot bars use published 0–100 scores, shrink to a 1.5rem minimum, then scroll horizontally.

Scores are published snapshots and are only comparable within the same benchmark, version, metric, and evaluation configuration.