Main Session
Sep 28
PQA 03 - Digital Health Innovation and Informatics, Patient Safety & Quality, and Radiation and Cancer Biology

2599 - Finding-Level Reliability of Vision-Language Models for Chest X-Ray Report Generation in Imaging-Based Clinical Decision Support

10:45am - 12:00pm ET
Poster Hall - Exhibit Hall A
Screen: 20
POSTER

Presenter(s)

Levent Sensoy, PhD - University of Miami/Jackson Health System, Miami, FL

L. Sensoy, and B. J. Rich; Department of Radiation Oncology, University of Miami / Sylvester Comprehensive Cancer Center, Miami, FL

Purpose/Objective(s):

Automated chest X-ray report generation using large vision-language models (VLMs) has gained increasing attention; however, rigorous finding-level evaluation across contemporary frontier and open-weight models remains limited. We systematically assessed the performance of state-of-the-art VLMs using a standardized benchmark and granular clinical claim–based evaluation framework. We hypothesized that domain-adapted VLMs would achieve higher finding-level precision and F1 scores than general-purpose models.

Materials/Methods:

Models were evaluated on the official MIMIC-CXR test split (2,461 reports; 285 patients). Reference reports created by radiologists served as ground truth. The reports generated were decomposed into atomic clinical claims representing discrete statements of presence, absence, or diagnostic uncertainty of radiologic findings. Trivial boilerplate normal statements were excluded unless contradicting documented abnormalities. Claims were matched to references to compute finding-level precision, recall, and F1. Evaluated systems included domain-adapted models (CheXOne; MedGemma-1.5-4B-IT; MedGemma-27B) and general-purpose VLMs (Qwen2.5-VL-3B; Qwen3-VL-8B; Haiku 4.5; Opus 4.5). To formally test the hypothesis, medical and generalist groups were compared using Welch’s t-test and Mann–Whitney U test with effect size estimation (Cohen’s d). Metrics are reported as mean ± standard deviation across cases.

Results:

Overall finding-level performance was modest across models. CheXOne achieved the highest precision (0.291 ± 0.261) and F1 (0.232 ± 0.219). MedGemma-1.5-4B-IT followed (F1 0.206 ± 0.200), whereas MedGemma-27B achieved F1 0.139 ± 0.178. Among general-purpose systems, F1 ranged from 0.091 ± 0.189 (Qwen2.5-VL-3B) to 0.153 ± 0.164 (Opus 4.5). Opus 4.5 demonstrated the highest recall (0.176 ± 0.195) and the greatest claim count per report (7.79 ± 2.96), reflecting increased verbosity without proportional precision gains.

Grouped analysis demonstrated significant superiority of domain-adapted models. Medical models achieved higher precision (0.263 vs 0.141; ?=0.121; 95% CI 0.111–0.131; Cohen’s d=0.51; p<0.001) and recall (0.203 vs 0.094; ?=0.109; 95% CI 0.101–0.117; d=0.59; p<0.001) compared with general-purpose models. Both parametric and non-parametric tests confirmed statistical significance.

Conclusion:

Current VLMs demonstrate limited reliability for clinically consistent chest X-ray report generation at the finding level. However, domain-adapted models significantly outperform general-purpose systems with moderate, practically meaningful effect sizes. These findings support the importance of medical specialization and rigorous claim-level evaluation before integration into imaging-based clinical decision support workflows.