GenAiHub
Task Planning & Knowledge Search
Open-ended research tasks · 2026

WideResearch

A broad research-agent benchmark for open-ended information gathering, synthesis, and answer construction across wide search spaces.

Metric

Score

Results shown

16

Publisher

BenchLM exact-source leaderboard

Snapshot

Fetched 2026-10-08T23:19:28.245Z

16 of 16 models

16 of 16 models

Higher is better

Me
Mercury 2
Hy
Hy4 preview
Qwen3.8 Max logo
Qwen3.8 Max
At
Atria Dawn Preview
Kimi K2.6 logo
Kimi K2.6
Or
Ornith-1.5-397B
do
dots3-note Preview
Claude Opus 4.5 logo
Claude Opus 4.5
Qwen3.6 Plus logo
Qwen3.6 Plus
Qwen3.5 397B logo
Qwen3.5 397B
Li
Ling 3.0 Flash
Kimi K2.5 logo
Kimi K2.5
GLM-5 logo
GLM-5
Or
Ornith-1.5-35B-A3B
Qwen3.6-35B-A3B logo
Qwen3.6-35B-A3B
Or
Ornith-1.5-9B
List bars visualize relative performance. Plot bars use published 0–100 scores, shrink to a 1.5rem minimum, then scroll horizontally.

Scores are published snapshots and are only comparable within the same benchmark, version, metric, and evaluation configuration.