2515 - Large Language Model-Based EMR Summarization as a Digital Safety Layer for Radiation Prescription
Presenter(s)
C. Lee1, J. H. Choi2, I. G. Hwang2, J. S. Kim1, and Y. K. Park3; 1Department of Radiation Oncology, Yonsei Cancer Center, Heavy Ion Therapy Research Institute, Yonsei University College of Medicine, Seoul, Korea, Republic of (South), 2Chung-Ang University College of Medicine, Seoul, Korea, Republic of (South), 3Department of Radiation Oncology, University of Texas Southwestern Medical Center, Dallas, TX
Purpose/Objective(s): Prescription errors in radiotherapy can result in significant patient harm despite routine peer review processes. Manual chart abstraction is time-consuming and may miss clinically meaningful deviations. We developed an automated framework using a large language model to extract structured clinical features from electronic medical records (EMR) and predict prescription parameters as a digital safety checkpoint.
Materials/Methods: Thoracic radiotherapy cases were retrospectively analyzed. In the first stage, unstructured EMR documentation was processed using a large language model (Llama-3.1-7B) to generate structured clinical feature summaries. These features included >20 parameters, such as treatment aim, tumor stage, concurrent chemotherapy, prior RT history, and others. In the second stage, we extracted tumor location and volume information from the treatment planning system. Finally, these structured features and the acquired planning parameters were used as inputs to random forest (RF) models to predict total dose, number of fractions, and biologically effective dose (BED). To isolate the impact of the summarization step, predictive performance of the RF models trained on AI-derived feature summaries was compared with performance of models trained on physician-abstracted structured summaries. Prediction accuracy, mean absolute error, and clinically significant deviation rates were evaluated. Subgroup analysis was performed for stereotactic body radiotherapy (SBRT) cases, defined as dose per fraction greater than 800 cGy and five or fewer fractions.
Results:
Random forest models trained on AI-derived feature summaries demonstrated performance comparable to physician-abstracted summaries for total dose and BED prediction (Table 1). R² values were 0.49 vs 0.55 for total dose and 0.64 vs 0.67 for BED (LLM vs physician). The largest discrepancy was observed in fraction number prediction, where physician-derived features substantially outperformed LLM-derived features (R² 0.78 vs 0.25). For SBRT classification, AI-derived features achieved slightly higher accuracy (0.925 vs 0.900).
Table 1. Comparison of random forest model performance using physician-abstracted versus LLM-derived feature summaries for prediction of total dose, fraction number, and biologically effective dose (BED).Conclusion: A large language model–assisted EMR summarization pipeline can extract clinically meaningful features that enable radiation prescription prediction approaching physician-derived abstraction. Although physician summaries showed superior fraction prediction, AI-derived features supported robust dose and BED modeling. This scalable approach may augment peer review and help reduce prescription errors in radiotherapy.
| w/ Human-summarized inputs | w/ LLM-summarized inputs | |||
| Prediction parameter | MAE | R2 | MAE | R2 |
| Total Dose (Gy) | 0.73 ± 0.61 | 0.55 | 0.71 ± 0.70 | 0.49 |
| Fractions | 1.74 ± 1.54 | 0.78 | 2.66 ± 3.42 | 0.25 |
| BED (Gy) | 1.61 ± 1.54 | 0.67 | 1.77 ± 1.49 | 0.64 |