evalgate
Eval-integrity statistics for AI benchmark claims — multiple-testing correction, power/MDE for model gaps, judge-bias a…
Live
Open / InstallLast updated July 27, 2026
Description
Eval-integrity statistics for AI benchmark claims — multiple-testing correction, power/MDE for model gaps, judge-bias and leaderboard-rank checks. Catches a benchmark number that won't survive a second look.
Author
ipezygj
Platform
mcp
Pricing model
free
Categories
mcp
Tags
hosting:remote-capable
mcp
glama
Capabilities
- MCP server