MedWER Introduces a Fixed-Term Evaluation Protocol for Medical Speech Recognition
A paper by Justin Behling posted to arXiv on 4 September 2026 introduces MedWER, a reproducible, model-free evaluation protocol and open-source tool for medical speech recognition. The metric uses a fixed, license-clean list of 19,373 drug, diagnosis, symptom, and injury-mechanism terms rather than an evaluation-time named-entity recognition model. Baselines for Moonshine base, Whisper, and MedASR on two open benchmarks include 95% confidence intervals.
A paper posted to arXiv on 4 September 2026 by Justin Behling introduces MedWER, a reproducible, model-free evaluation protocol and open-source tool for medical speech recognition. The paper is titled MedWER: A Reproducible, Model-Free Evaluation Protocol for Medical Speech Recognition and appears as arXiv 2609.05728, with version 1 submitted Friday, 4 September 2026 at 21:19:18 UTC. MedWER is both an evaluation protocol and an open-source tool aimed at medical automatic speech recognition. Overall word error rate hides clinically critical errors. The paper states that a transcript can be 95% correct and still swap one drug for another. The usual fix weights errors on medical entities. That usual fix almost always depends on an evaluation-time named-entity recognition model or cloud API. Dependence on that model or API makes the metric's denominator a versioned black box. MedWER sets the denominator to a fixed, license-clean term list. The list contains 19,373 drug, diagnosis, symptom, and injury-mechanism entries projected from public sources. The protocol couples a pinned text normalizer with a phrase-aware term-restricted WER, called the MedWER. The only versioned component is a normalizer dependency held at an exact release. That normalizer is checked against committed golden fixtures. Coverage is validated against an independent provincial drug-benefit file the list was not built from. The matching heuristic is calibrated against ground-truth entity spans. Baselines for Moonshine base, Whisper, and MedASR on two open benchmarks are scored with the released tool. Those baselines are reported with 95% confidence intervals from resampled per-utterance scores. Scoring with the released tool uses the fixed term list and the pinned normalizer rather than an evaluation-time named-entity recognition model or cloud API.