This skill should be used when the user asks to "evaluate agent performance", "build test framework", "measure agent qu…
This skill should be used when the user asks to "evaluate agent performance", "build test framework", "measure agent quality", "create evaluation rubrics", or mentions LLM-as-judge, multi-dimensional evaluation, agent testing, or quality gates for agent pipelines.
muratcankoylan
cli
free
Others in the same category, ranked by how often they are opened.