Reporting Practice Matters: The Impact of Reference Choice on Chest X-ray Report Evaluation
Daniel P. Jeong, Charles Q. Li, Hossein Hosseiny, Nitya M. Bhalla, Fatma Uyar Morency, Pradeep Ravikumar, Zachary C. Lipton, Michael Oberst
- Published
- Sep 16, 2026 — 17:28 UTC
Problem
Variations in radiology reporting practices significantly influence the evaluation of AI-based radiology report generation models. This paper addresses the gap in understanding how these variations affect model performance, particularly in the context of chest X-ray reports. The work is presented as a preprint and has not undergone peer review.
Method
The authors propose a radiologist-informed taxonomy of variations in radiology reporting practices. This taxonomy serves as a framework for analyzing how different reporting styles can impact the evaluation of AI models. The core technical contribution is the method called ReRef, which rewrites reference reports according to the established taxonomy while maintaining the clinical interpretation of the reports. The dataset used for evaluation is MIMIC-CXR-Ext-ReRef, which consists of 120 pairs of original and alternative reference reports derived from the MIMIC-CXR dataset. This structured approach allows for a systematic examination of how variations in reporting can influence model rankings and evaluations.
Results
The evaluation of AI models using the ReRef method revealed notable changes in model rankings when reference reports were condensed. Specifically, the Libra model dropped from first to second place, while the CheXOne model improved from third to first place when assessed against the RadCliQ-v1 benchmark. These results underscore the impact of reference choice on model performance and highlight the importance of standardized reporting practices in the evaluation of AI systems.
Limitations
The authors acknowledge that current evaluation metrics may not adequately separate clinical interpretation from adherence to specific reporting practices. This limitation suggests that while the proposed taxonomy and ReRef method provide valuable insights, further refinement of evaluation metrics is necessary to ensure that they accurately reflect model performance independent of reporting style. Additionally, the study is limited by the scope of the MIMIC-CXR dataset, which may not encompass the full diversity of radiology reporting practices.
Why it matters
This work has significant implications for the development and evaluation of AI models in radiology. By highlighting the influence of reporting variations on model performance, it calls for a reevaluation of how AI systems are assessed in clinical settings. The introduction of a structured taxonomy for reporting practices can guide future research in creating more robust evaluation frameworks, ultimately leading to improved AI systems that are better aligned with clinical needs and practices.
By Callan Zhang · Sep 16, 2026 · Editorial standards →
Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.
Source: arXiv cs.AI
