Exposes a run_suite tool to evaluate whether an AI agent is safe to operate internal web apps, scoring task completion …
Exposes a run_suite tool to evaluate whether an AI agent is safe to operate internal web apps, scoring task completion and forbidden-action violations to gate CI/CD pipelines.
RudrenduPaul
mcp
free
No benchmark results have been added yet.