Main Session
Sep 28
PQA 03 - Digital Health Innovation and Informatics, Patient Safety & Quality, and Radiation and Cancer Biology

2456 - Leveraging Large Language Models and Text Embeddings on Radiology Reports to Predict Radiation Therapy Treatment Outcomes in Hepatocellular Carcinoma

10:45am - 12:00pm ET
Poster Hall - Exhibit Hall A
Screen: 13
POSTER

Presenter(s)

Jie Fu, PhD Headshot
Jie Fu, PhD - University of Washington, Seattle, WA

J. Fu1, M. J. Nyflot1, Y. Kim1, Y. He1, S. R. Bowen1, C. Grassberger1, S. Apisarnthanarax2, and S. Cui1; 1Department of Radiation Oncology, University of Washington/Fred Hutchinson Cancer Center, Seattle, WA, 2Department of Radiation Oncology, University of Washington, Seattle, WA

Purpose/Objective(s): Diagnostic radiological reports contain rich, detailed information regarding tumor characteristics and disease burden that are often underutilized in conventional predictive models relying on structured data. This study aims to explore the efficacy of leveraging large language models (LLMs) and domain-specific text embeddings to encode unstructured text extracted from radiological reports for predicting clinical outcomes, including Child-Pugh (CP) score change and overall survival (OS) in hepatocellular carcinoma (HCC) patients undergoing radiation therapy (RT).

Materials/Methods: A retrospective analysis was conducted on 119 patients with diagnostic MRI and/or CT within 3 months prior to RT. Clinical endpoints included OS and CP score increase (=2 points, CP2+) within 6 months post-RT, indicating liver function decline. Text from “Impression” and “Findings” sections of radiology reports was extracted using optical character recognition tools, followed by denoising and manual review for data fidelity and de-identification. For patients with two pre-RT scans, report texts were concatenated into a single patient-level document for modeling. We compared two approaches: (1) zero-shot prediction using a PHI-compliant GPT model (GPT-5.0) via Azure API with customized prompts tailored to clinical endpoints, and (2) text embeddings from Bio-Clinical BERT were input into gradient boosting (XGBoost) classifiers for CP2+ prediction and XGBoost Cox models for OS prediction. Performance was assessed by the area under the curve (AUC) for CP2+ and C-index for OS, with 5-fold cross-validation.

Results: Embedding-based models demonstrated robust performance with a CP2+ AUC of 0.70 (95% CI: 0.56-0.83) and an OS C-index of 0.64 (95% CI: 0.58-0.70). Zero-shot GPT predictions yielded moderate performance (AUC: 0.65 [95% CI: 0.51-0.77]; C-index: 0.61 [95% CI: 0.54-0.67]), without statistically significant differences relative to embedding-based methods. Subgroup analysis revealed that zero-shot GPT performs worse in predicting CP2+ among CP-B/C class patients compared to embedding-based methods, with an AUC of 0.36 (95% CI: 0.14-0.60) versus 0.70 (95% CI: 0.51-0.89) respectively (p=0.02). Embedding-based models showed more consistent performance in both CP-A and CP-B/C patients across the two prediction tasks.

Conclusion: The integration of LLM and text embeddings from unstructured radiological reports offers a promising, yet underutilized, avenue for enhancing prognostic modeling in HCC patients undergoing radiation therapy. Domain-specific training of embedding models demonstrated incremental gains in performance compared to zero-shot LLM-generated predictions. Future research focusing on fine-tuning domain-adapted models, integrating multimodal data including imaging, and implementing advanced prompt engineering strategies for LLMs holds potential to further enhance predictive accuracy and support personalized treatment stratification.