Document-level MT evaluation protocols fail to capture cross-segment consistency
Ahrii Kim reports on arXiv that document-level machine translation evaluation does not elicit document-level judgments. A MIX condition that combined segments from different systems produced statistically equivalent scores, rankings, and error annotations versus coherent documents across 18,420 expert English-to-Korean labels and 14 automatic metrics. Raters identified coherent passages as a single translator's work in 87.3 percent of trials. The protocol, not the annotator, is blind.
A paper posted to arXiv on 5 September 2026 finds that document-level machine translation evaluation does not measure the document-level quality it is built to capture. Ahrii Kim reports that expert annotators produce statistically equivalent scores, system rankings, and error annotations whether they judge coherent documents or incoherent documents whose segments come from different systems. The same equivalence holds across 14 automatic metrics. The result is drawn from 18,420 expert English-to-Korean annotations. Kim concludes that what is blind is the protocol, not the annotator, and that resources invested in document-level systems, metrics, and annotation may not be measuring what they are intended to measure. Document-level machine translation evaluation extends segment-level protocols by presenting full documents to annotators, on the assumption that such presentation elicits document-level judgments. The paper, titled The Blindness of Document-Level Translation Evaluation, tests that assumption with a counterfactual condition called MIX. In the MIX condition, each document combines segments drawn from different systems. Document-level presentation is preserved, but cross-segment consistency is broken. Annotators still see a full document. The document is no longer the output of a single system. Scores, system rankings, and error annotations remained statistically equivalent between coherent documents and MIX documents for both the human annotators and the 14 automatic metrics. Perception does not explain this equivalence. Shown matched passages, raters identified the coherent one as the work of a single translator in 87.3 percent of trials. Annotators can tell a coherent document from a mixed one when asked. Document presentation does change how annotators work, the paper states, but that change does not reach the recorded output. The recorded scores, rankings, and error annotations do not reflect the document-level property the protocol is meant to capture. The concern is not that scores fall short, but that the resources invested in document-level systems, metrics, and annotation may not be measuring what they are intended to measure. The paper was submitted on 5 September 2026.