Task Planning & Knowledge Search
Open-ended research tasks · 2026WideResearch
A broad research-agent benchmark for open-ended information gathering, synthesis, and answer construction across wide search spaces.
Metric
Score
Results shown
16
Publisher
BenchLM exact-source leaderboard
Snapshot
Fetched 2026-10-08T23:19:28.245Z
16 of 16 models
16 of 16 models
Higher is better
Me
Mercury 2Hy
Hy4 previewAt
Atria Dawn PreviewOr
Ornith-1.5-397Bdo
dots3-note PreviewLi
Ling 3.0 FlashOr
Ornith-1.5-35B-A3BOr
Ornith-1.5-9BList bars visualize relative performance. Plot bars use published 0–100 scores, shrink to a 1.5rem minimum, then scroll horizontally.
Scores are published snapshots and are only comparable within the same benchmark, version, metric, and evaluation configuration.