Code & Software Engineering
1,865 repository problems · 2025SWE-bench Pro
A long-horizon repository benchmark built to test realistic software engineering work. Its scores need a task-quality and setup check before they support a coding-agent decision.
Metric
Resolved
Results shown
78
Publisher
BenchLM exact-source leaderboard
Snapshot
Fetched 2026-10-08T23:19:28.332Z
78 of 78 models
78 of 78 models
Higher is better
Sa
Sakana Fugu-UltraHy
Hy4 previewBe
BeamOr
Ornith-1.5-397BOr
Ornith-1.0-397Bdo
dots3-note PreviewAt
Atria Dawn PreviewOr
Ornith-1.5-35B-A3BLa
Laguna S 2.1Sa
Sakana FuguLi
Ling 3.0 FlashSt
Step 3.7 FlashIn
Inkling-SmallIn
InklingMA
MAI-Thinking-1Or
Ornith-1.0-35BLa
Laguna M.1La
Laguna XS 2.1Or
Ornith-1.5-9BLa
Laguna XS.2Or
Ornith-1.0-9BLo
LongCat-Flash-Lite-SparseGr
Granite 4.2 30BLL
LLaDA2.2-flashMe
Mellum2.1-12B-A2.5B-ThinkingGr
Granite 4.2 8BMi
MiniCPM5-2BList bars visualize relative performance. Plot bars use published 0–100 scores, shrink to a 1.5rem minimum, then scroll horizontally.
Scores are published snapshots and are only comparable within the same benchmark, version, metric, and evaluation configuration.