Use this skill when the user wants verifiable reasoning tasks to benchmark or test an LLM or agent — reproducible puzzl…
Use this skill when the user wants verifiable reasoning tasks to benchmark or test an LLM or agent — reproducible puzzle task sets (sudoku, word search) with objective, by-construction grading. No answer key to trust: answers are verified against the rules and the grid.
caravaca-labs
cli
free
Others in the same category, ranked by how often they are opened.