Main Session
Sep 28
PQA 03 - Digital Health Innovation and Informatics, Patient Safety & Quality, and Radiation and Cancer Biology

2575 - BrainMATTER: Brain Metastasis Auto-Tracking through Text Extraction from Reports

10:45am - 12:00pm ET
Poster Hall - Exhibit Hall A
Screen: 19
POSTER

Presenter(s)

Mitchell Parker, MD, PhD Headshot
Mitchell Parker, MD, PhD - Memorial Sloan Kettering Cancer Center, New York, NY

M. I. Parker1, J. T. Rosenthal2, R. R. Patel3, E. Miao3, A. Skakodub4, E. Wo5, Z. Yazdani6, D. T. Bergman7, C. Kinslow3, Y. Yu3, B. S. Imber8, G. Cederquist3, C. B. Jackson3, M. Zinovoy3, C. Fong8, M. Waters8, J. Jee8, A. Li9, L. Z. Braunstein8, and L. R. G. Pike8; 1Drexel University College of Medicine, Philadelphia, PA, 2Weill Cornell Medical College, New York, NY, 3Department of Radiation Oncology, Memorial Sloan Kettering Cancer Center, New York, NY, 4Department of Radiation Oncology, Icahn School of Medicine at Mount Sinai, New York, NY, 5Memorial Sloan Kettering Cancer Center, New York City, NY, 6Rutgers Robert Wood Johnson Medical School, New Brunswick, NJ, 7Geisel School of Medicine at Dartmouth, Hanover, NH, 8Memorial Sloan Kettering Cancer Center, New York, NY, 9Department of Medical Physics, Memorial Sloan Kettering Cancer Center, New York, NY

Purpose/Objective(s):

Neuroimaging-derived clinical characteristics are critical for risk stratification in patients with brain metastases (BM) but are often documented as narrative text in radiology reports, limiting scalable extraction. We hypothesized that large language model (LLM) pipelines can accurately abstract BM-specific imaging variables from unstructured brain MRI reports with fidelity comparable to or exceeding manual abstraction.

Materials/Methods:

We curated brain MRI reports from an IRB-approved cohort of patients with BM at a single institution (2010–2024). Two LLM pipelines were developed: (1) a report-level module to extract BM count and largest BM dimension from the diagnostic brain MRI report; and (2) a longitudinal impression-level module to identify BM diagnosis date from concatenated impressions. Analyses used GPT-5.1 via an institutionally approved, HIPAA-compliant service. LLM outputs were benchmarked against a gold-standard dataset curated by physician abstractors. For BM count, discordant cases underwent adjudication by a medical student and resident to determine correctness. Cases were excluded if the LLM refrained due to missing or ambiguous content.

Results:

The cohort included 863 patients (median follow-up 25.1 months; IQR 9.1–51.6). Median age at BM diagnosis was 61 years (IQR 53–70), and 70.5% were female. Patients had a median of 8 brain MRI reports (IQR 5–14); BM diagnosis occurred at the first report for 586 patients (67.9%). Median largest BM dimension was 22 mm (range 2–70). BM count distribution was: 1 (27.6%), 2 (12.7%), 3 (7.0%), 4 (3.0%), 5–9 (10.9%), 10+ (7.8%), Unknown (31.1%). For BM count (n=520), initial LLM-versus-human accuracy was 85.6% (Cohen's ?=0.81). Among 75 discordant cases, adjudication determined the LLM was correct in 65 (86.7%), the human abstractor in 6 (8.0%), and neither in 4 (5.3%), yielding post-adjudication LLM accuracy of 98.1%. For largest lesion dimension (n=549), the LLM achieved 83.2% exact agreement and 92.9% accuracy within +/- 3 mm versus non-adjudicated human abstraction (Spearman r=0.96). For BM diagnosis date (n=841), LLM accuracy was 90.6% versus the non-adjudicated human gold standard (mean absolute error 19 days).

Conclusion:

LLM accuracy for BM count exceeded 98% after adjudication, with correct predictions in over 86% of initially discordant cases. Non-adjudicated accuracy exceeded 90% for diagnosis date and largest lesion dimension within +/- 3 mm. LLM-based pipelines enable rapid, high-throughput abstraction with excellent accuracy, and adjudication of discordant cases may further improve performance. This approach can substantially accelerate the creation of research-ready datasets for outcomes modeling. Future work will assess cross-institution generalizability and expand extraction to more complex variables.