GenAiHub
AI Research

UniRRM and MixReward Unify Multilingual Reasoning Rewards Across Evaluation Paradigms

1 min read · 270 wordsAI-curated · Powered by AtmezAI
UniRRM and MixReward Unify Multilingual Reasoning Rewards Across Evaluation Paradigms

A 5 September 2026 arXiv paper by Lai Peng introduces UniRRM, a unified reasoning reward model, and MixReward, a dataset covering six domains and 103 languages with pairwise and listwise examples. UniRRM uses a staged reasoning chain to generate task-generic and instruction-specific criteria. UniRRM-8B and UniRRM-14B approach state-of-the-art results for their size, transfer to unseen evaluation paradigms, and are supported by ablation studies.

Researchers led by Lai Peng have introduced UniRRM, a unified reasoning reward model that supports multiple languages and evaluation paradigms, together with MixReward, a large-scale multilingual dataset. The paper appeared on arXiv on 5 September 2026. Reinforcement learning excels on tasks with verifiable rewards, but in open-ended tasks the reliability of reward models remains a key challenge. Existing solutions either depend on costly proprietary LLM-as-a-Judge systems or opaque scalar reward models that lack interpretability. Recent works on generative reward models offer a promising alternative, but they remain constrained by static evaluation criteria, fragmented evaluation paradigms, and limited multilingual support. To address these challenges, the team introduces MixReward, spanning six domains and 103 languages and containing both pairwise and listwise data, and proposes UniRRM. UniRRM uses a staged reasoning chain to dynamically generate task-generic and instruction-specific criteria, enabling fine-grained, input-adaptive judgments while maintaining consistency across languages. MixReward’s six-domain, 103-language coverage and dual pairwise-listwise structure give the reward model training signal across the settings that currently fragment evaluation. The staged chain first forms general criteria and then instruction-specific ones before producing a judgment, rather than applying a fixed rubric or a single scalar. Experiments demonstrate that UniRRM-8B and UniRRM-14B achieve performance close to the state-of-the-art for models of comparable size across multiple benchmarks, and are effective for unseen evaluation paradigms. Ablation studies validate the reliability and effectiveness of UniRRM. For practitioners scoring open-ended model outputs, the 8B and 14B models offer near state-of-the-art quality at those sizes, multilingual operation over 103 languages, and explicit generated criteria instead of an uninterpretable number, including on evaluation paradigms the model did not see during training.

Verified sources · 1