Main Session
Sep 29
SS 44 - Using AI and Other Software to Elevate Patient Safety and Quality

336 - Teaching Language Models Safe Clinical Decision Making: Risk-Informed Training to Reduce Catastrophic Errors in Oncology

02:25pm - 02:35pm ET
Room 256

Presenter(s)

Reza Khanmohammadi, MS - Michigan State University, East Lansing, MI

R. Khanmohammadi1, A. I. Ghanem2, Y. Jee2, A. R. Bhatnagar2, S. Siddiqui2, M. A. Elshaikh2, H. Bagher-Ebadian2, B. Movsas2, M. M. Ghassemi1, and K. Thind2; 1Department of Computer Science and Engineering, Michigan State University, East Lansing, MI, 2Department of Radiation Oncology, Henry Ford Health, Detroit, MI

Purpose/Objective(s): Language models (LLMs) are increasingly explored for oncology decision support, including radiation therapy planning and EHR-based chart review. However, standard training treats all errors equally, ranging from minor mislabeling to major adverse events. We hypothesized that clinician-validated risk scores during fine-tuning would teach models to avoid catastrophic errors and shift failures toward safer outcomes.

Materials/Methods: We fine-tuned three open-source LLMs (1.5B–3B parameters) on EHRNoteQA (1,659 oncology EHR-derived QA samples). Evaluation used three test sets from CORAL, a publicly available oncology QA benchmark published in the New England Journal of Medicine AI, spanning increasing clinical complexity: Contextual (1,126 retrieval-based questions), Synthesis (763 multi-hop reasoning), and Clinical Inference (377 expert-level diagnostic reasoning). A frontier LLM assigned clinical risk score to each incorrect answer: 1 (minor, no patient harm), 5 (suboptimal management), or 10 (life-threatening error). Two physicians independently validated risk labels on a stratified sample (n=300 questions, 900 option-level judgments; Cohen’s ?=0.51–0.64; within-one-category agreement 93–96%). Training combined cross-entropy with a risk-weighted penalty that increases loss when models confidently predict dangerous outcomes. Changes in safety-adjusted accuracy (Risk-Weighted Accuracy, RWA) and average harm severity per error (clinical error cost) were analyzed.

Results: Risk-informed training improved safety-adjusted accuracy across all datasets and models, with the largest gain on expert-level Clinical Inference questions: up to 7.2% RWA improvement and 14.7% reduction in average error severity. When these models erred, they were more likely to select minor rather than dangerous mistakes. The model confidence also improved, enabling clinicians to better flag uncertain cases for review (51–70% calibration improvement across all conditions).

Conclusion: Risk-informed training reduced dangerous errors in oncology LLMs by penalizing high-risk mistakes during fine-tuning. When errors occurred, LLMs shifted towards clinically minor consequences, and model confidence became more reliable for flagging cases needing human oversight. Therefore, risk-informed training should be implemented in radiation oncology workflow to prioritize safe failures over raw accuracy and ensure clinically safe LLM deployment.

Abstract 336 - Table 1: Safety-Adjusted Accuracy and Error Severity After Risk-Informed Training

Model

Dataset

RWA

?RWA (%)

?Error Severity (%)

1.5B

Contextual

0.80

+1.5

+6.4

1.5B

Clinical Inference

0.73

+5.9

+13.2

3B-A

Synthesis

0.82

+2.4

+9.7

3B-A

Clinical Inference

0.72

+1.1

+2.7

3B-B

Contextual

0.79

+1.6

+5.7

3B-B

Clinical Inference

0.72

+7.2

+14.7

OVERALL

All

Median

0.76

+2.0

+8.05

IQR

0.72-0.80

1.5-5.0

5.9-12.3