A benchmark for web-browsing agents that must search, inspect sources, gather evidence, and return the correct answer to research-oriented questions.
Metric
Accuracy
Results shown
40
Publisher
BenchLM exact-source leaderboard
Snapshot
Fetched 2026-08-22T12:55:14.749Z
40 of 40 models
Higher is better
Scores are published snapshots and are only comparable within the same benchmark, version, metric, and evaluation configuration.