A long-horizon repository benchmark built to test realistic software engineering work. Its scores need a task-quality and setup check before they support a coding-agent decision.
Metric
Resolved
Results shown
63
Publisher
BenchLM exact-source leaderboard
Snapshot
Fetched 2026-08-22T12:55:14.783Z
63 of 63 models
Higher is better
Scores are published snapshots and are only comparable within the same benchmark, version, metric, and evaluation configuration.