A broad research-agent benchmark for open-ended information gathering, synthesis, and answer construction across wide search spaces.
Metric
Score
Results shown
14
Publisher
BenchLM exact-source leaderboard
Snapshot
Fetched 2026-08-22T12:55:14.716Z
14 of 14 models
Higher is better
Scores are published snapshots and are only comparable within the same benchmark, version, metric, and evaluation configuration.