Task Planning & Knowledge Search
Research questions requiring browsing · 2025BrowseComp
A benchmark for web-browsing agents that must search, inspect sources, gather evidence, and return the correct answer to research-oriented questions.
Metric
Accuracy
Results shown
47
Publisher
BenchLM exact-source leaderboard
Snapshot
Fetched 2026-10-08T23:19:28.309Z
47 of 47 models
47 of 47 models
Higher is better
At
Atria Dawn PreviewSt
Step 5 PreviewOr
Ornith-1.5-397Bdo
dots3-note PreviewIn
Inkling-SmallBe
BeamIn
InklingSt
Step 3.7 FlashAg
Agents-A1Li
Ling 3.0 FlashOr
Ornith-1.5-35B-A3BAg
Agents-A1-4BOr
Ornith-1.5-9BSo
Solar Pro 4Lo
LongCat-Flash-Lite-SparseNe
Nemotron 3 UltraNe
Nemotron 3.5 Lightning 30B A3B NVFP4List bars visualize relative performance. Plot bars use published 0–100 scores, shrink to a 1.5rem minimum, then scroll horizontally.
Scores are published snapshots and are only comparable within the same benchmark, version, metric, and evaluation configuration.