Main Session
Sep 28
PQA 03 - Digital Health Innovation and Informatics, Patient Safety & Quality, and Radiation and Cancer Biology

2628 - Evaluating Multimodal LLMs for Information Extraction from Oncology Reports Requires a Clinically Curated Ground Truth: Two-Phase Evaluation of GPT-4.1 vs GPT-4.0

10:45am - 12:00pm ET
Poster Hall - Exhibit Hall A
Screen: 21
POSTER

Presenter(s)

Nikhil Thaker, MD, MBA, MHA - Capital Health Medical Center Hopewell, Pennington, NJ

A. Loaiza-Bonilla1, N. G. Thaker2, E. Tuysuz3, S. Yerdan3, R. Kraft2, and S. Kurnaz3; 1St Luke’s University Health Network, Easton, PA, 2Capital Health, Pennington, NJ, 3MassiveBio, Easton, PA

Purpose/Objective(s):

Evaluate the ability of large language models (LLMs) to extract clinically relevant imaging findings from oncological records, and quantify how quality of the reference standard (uncurated vs expert-curated “gold standard”) affects measured performance. We hypothesized that (1) expert-curated reference summaries would substantially improve all performance metrics compared with heterogeneous baseline references, and (2) a multimodal model (GPT-4.1) would outperform a text-only model (GPT-4.0), particularly when provided with image-based input and full-record context.

Materials/Methods:

We performed a 2-phase evaluation on 40 deidentified oncology records from a single source system. In Phase 1, model outputs were compared against existing, uncurated reference summaries. In Phase 2, 20 randomly selected records were reannotated by a board-certified radiologist to create structured, “gold standard” summaries (one line per imaging study, including modality and key findings). 16 configurations were tested, combining 2 model versions (GPT-4.0 vs GPT-4.1), 2 prompts (concise vs detailed), 2 input modalities (OCR text vs document images; image only for GPT-4.1), and 2 document scopes (tagged imaging pages vs whole record). Outputs were compared with references using lexical metrics (ROUGE-1/L, BLEU) and a semantic alignment metric (Kullback–Leibler [KL] divergence).

Results:

Against uncurated references (Phase 1), performance appeared modest across configurations (best ROUGE-1 ˜0.45, BLEU ˜0.15, KL ˜7.73), with substantial apparent semantic divergence. When evaluated against expert-curated gold-standard summaries (Phase 2), measured performance improved across nearly all settings. The top configuration—GPT-4.1 with image input, whole-record scope, and concise prompt—achieved ROUGE-1 0.57, ROUGE-L 0.55, BLEU 0.25, and KL 5.96. Under gold-standard conditions, GPT-4.1 consistently outperformed GPT-4.0 across lexical and semantic metrics; image-based input outperformed text-only input for GPT-4.1; whole-record input outperformed pre-tagged pages; and the concise prompt outperformed the more detailed prompt. Qualitative review showed that high-scoring outputs closely mirrored expert summaries of imaging findings.

Conclusion:

The quality of the reference standard impacts LLM performance in extraction of clinical information, with expert-curated summaries outperforming uncurated benchmarks, and multimodal GPT-4.1 outperforms the text only GPT-4.0. These findings underscore the need to co-develop clinical AI systems and their evaluation frameworks, investing in high-fidelity benchmarks so that model assessment reflects a clinically meaningful “ground truth.”