GenAiHub
Code & Software Engineering
Epoch live leaderboard

SWE-bench Verified

Human-validated software engineering issues from real-world Python repositories.

Metric

Accuracy

Results shown

33

Publisher

Epoch AI live leaderboard

Snapshot

Fetched 2026-10-08T23:19:31.697Z

33 of 33 models

33 of 33 models

Higher is better

Claude Opus 4.7 (max) logo
Claude Opus 4.7 (max)
GPT-5.5 (xhigh) logo
GPT-5.5 (xhigh)
Gemini 3.5 Flash (high) logo
Gemini 3.5 Flash (high)
Claude Opus 4.6 (no thinking) logo
Claude Opus 4.6 (no thinking)
GLM-5.2 (max) logo
GLM-5.2 (max)
DeepSeek v4 (max) logo
DeepSeek v4 (max)
Qwen3.7 Max logo
Qwen3.7 Max
GPT-5.4 (high) logo
GPT-5.4 (high)
Qwen3.6 Max Preview (thinking) logo
Qwen3.6 Max Preview (thinking)
Kimi K2.6 logo
Kimi K2.6
Claude Opus 4.5 (no thinking) logo
Claude Opus 4.5 (no thinking)
Gemini 3.1 Pro Preview (Custom tools) logo
Gemini 3.1 Pro Preview (Custom tools)
Gemini 3 Flash Preview logo
Gemini 3 Flash Preview
Claude Sonnet 4.6 (no thinking) logo
Claude Sonnet 4.6 (no thinking)
GPT-5.3 Codex (high) logo
GPT-5.3 Codex (high)
GLM-5.1 logo
GLM-5.1
Kimi K2.5 logo
Kimi K2.5
GPT-5.2 (high) logo
GPT-5.2 (high)
GPT-5 (high) logo
GPT-5 (high)
Claude Opus 4.1 logo
Claude Opus 4.1
Gemini 3 Pro Preview logo
Gemini 3 Pro Preview
GLM-5 logo
GLM-5
GPT-5 (medium) logo
GPT-5 (medium)
Claude Sonnet 4.5 (no thinking) logo
Claude Sonnet 4.5 (no thinking)
Claude Opus 4 logo
Claude Opus 4
GPT-5.1 (high) logo
GPT-5.1 (high)
GPT-5 mini (medium) logo
GPT-5 mini (medium)
o3 (medium) logo
o3 (medium)
Claude 3.7 Sonnet logo
Claude 3.7 Sonnet
Qwen3.6 Plus (thinking) logo
Qwen3.6 Plus (thinking)
Gemini 2.5 Pro logo
Gemini 2.5 Pro
GPT-4.1 logo
GPT-4.1
GPT-4o (Nov 2024) logo
GPT-4o (Nov 2024)
List bars visualize relative performance. Plot bars use published 0–100 scores, shrink to a 1.5rem minimum, then scroll horizontally.

Scores are published snapshots and are only comparable within the same benchmark, version, metric, and evaluation configuration.