GenAiHub
AI Research

AhaBench Tests Whether Language Agents Learn From Prior Experience

2 min read · 313 wordsAI-curated · Powered by AtmezAI
AhaBench Tests Whether Language Agents Learn From Prior Experience

AhaBench, submitted to arXiv on 30 June 2026, tests whether a fixed language model improves later behavior after useful experience when support is removed, changed, or delayed. The suite covers puzzles, Project-Euler-style math, and a vending simulator with Initial Score, Post-Experience Score, and Learning Lift. Claude Opus 4.6 led Post-Experience Score at 64.3 and Learning Lift at +25.8, with Gemini 3.1 Pro at 63.4. Tasks, validators, and simulator code are released.

A paper submitted to arXiv on 30 June 2026 introduces AhaBench, a benchmark that asks whether language agents learn from prior experience. AhaBench poses an operational question: when a fixed model receives useful experience, does its later behavior improve under a related evaluation condition where the obvious support has been removed, changed, or delayed? The work, from Zerui Cheng, notes that modern language agents are expected to operate over long horizons: they ask follow-up questions, reuse worked examples, handle tool feedback, and adapt to delayed consequences. Most evaluations still reset the agent after a prompt or score only the final state of one trajectory. The suite contains three components. Aha-Puzzle tests no-hint exploration after solved hidden-state puzzles. Aha-Euler turns Project-Euler-style mathematical ideas into generated taught and held-out tasks with exact validators. Aha-Vending, an open-source implementation inspired by Vending-Bench, tests whether a simulated vending agent remains profitable while handling delayed feedback and operational incidents. AhaBench reports a three-part scorecard. Initial Score measures starting competence. Post-Experience Score measures the later empirical outcome. Learning Lift is their difference. This decomposition is the main empirical message: models that use visible support well, models that reach high post-experience scores, and models that improve most during a run are not always the same. On the common eight-model panel, Claude Opus 4.6 leads aggregate Post-Experience Score at 64.3 and aggregate Learning Lift at +25.8, with Gemini 3.1 Pro close behind at 63.4. The component results explain the split: puzzle traces raise supported scores but often fail to become no-hint exploration behavior; Aha-Euler full teaching reaches 78.6-100.0% while answer-only transfer ranges from 0.0 to 73.9%; and Aha-Vending separates profitable incident handling from bankruptcy and no-order failure. The authors release benchmark tasks, rubrics, validators, simulator code, and interfaces for evaluating new agents. Evaluators can compare Initial Score, Post-Experience Score, and Learning Lift instead of scoring only the final state of one trajectory.

Verified sources · 1