Main Session
Sep 28
PQA 03 - Digital Health Innovation and Informatics, Patient Safety & Quality, and Radiation and Cancer Biology

2502 - Risk-Controlled Selective Automation of Oncology Chart Review via Representational Stability Calibration

10:45am - 12:00pm ET
Poster Hall - Exhibit Hall A
Screen: 30
POSTER

Presenter(s)

Reza Khanmohammadi, MS - Michigan State University, East Lansing, MI

R. Khanmohammadi1, A. I. Ghanem2, S. Siddiqui2, H. Bagher-Ebadian2, B. Movsas2, M. M. Ghassemi1, and K. Thind2; 1Department of Computer Science and Engineering, Michigan State University, East Lansing, MI, 2Department of Radiation Oncology, Henry Ford Health, Detroit, MI

Purpose/Objective(s): Radiation oncology workflows generate extensive documentation across consult intake, on-treatment visits, and follow-up, adding to the clinician burden. Locally deployed language models could assist with routine chart review, but probability-based confidence fails to detect clinically important reasoning errors. We hypothesized that internal representational stability, how consistently a model’s hidden states respond to small input perturbations, provides a more reliable safety signal, enabling selective automation where uncertain cases are deferred to clinicians.

Materials/Methods: We developed Calibrating Confidence via Perturbation Stability (CCPS), which evaluates the stability of a model’s final hidden-layer representation across targeted perturbations of the same clinical note and question pair. Stable cases are automated; unstable cases are deferred for human review. Thresholds were tuned on a held-out training split (n=1,659) from an electronic health record QA corpus. CCPS was evaluated on an independent, multi-institutional test set of expert-labeled breast and pancreatic cancer notes (n=2,266) across three complexity tiers: Contextual (n=1,126), Synthesis (n=763), and Clinical Inference (n=377). Five open-weight, locally deployable models (0.5B to 3B parameters) were compared against probability-thresholding (PT) and internal-knowledge probing (PIK) baselines. Safety was defined as no greater than 5% incorrect answers among cases that the model automated rather than deferred.

Results: CCPS automated a larger share of cases than either baseline while keeping errors among non-deferred cases within the 5% threshold (Table 1). The next-best method (PIK) achieved lower yields of 32%, 13%, and 44% under the same constraint; PT failed to automate any cases in two tiers. CCPS produced the highest safe yield for every individual model in the Clinical Inference tier, and its advantage over baselines widened with task complexity.

Conclusion: Representational-stability calibration provides a principled safety layer for deploying small, locally hosted language models in oncology documentation review. By deferring uncertain cases rather than forcing predictions, this approach maintains strict error limits while reducing clinician review volume, supporting prospective evaluation in radiation oncology workflows.

Table 1. CCPS case-level triage across five locally hostable models. Automated cases are answered without human review; error rate reflects the proportion answered incorrectly. Median values; safety threshold: =5% error.

Complexity Tier

Total

Cases

Automated by CCPS

Deferred to Clinician

Error Among Automated

Contextual

1,126

405 (36%)

721 (64%)

4%

Synthesis

763

114 (15%)

649 (85%)

4%

Clinical Inference

377

189 (50%)

188 (50%)

2%