2502 - Risk-Controlled Selective Automation of Oncology Chart Review via Representational Stability Calibration
Presenter(s)
R. Khanmohammadi1, A. I. Ghanem2, S. Siddiqui2, H. Bagher-Ebadian2, B. Movsas2, M. M. Ghassemi1, and K. Thind2; 1Department of Computer Science and Engineering, Michigan State University, East Lansing, MI, 2Department of Radiation Oncology, Henry Ford Health, Detroit, MI
Purpose/Objective(s): Radiation oncology workflows generate extensive documentation across consult intake, on-treatment visits, and follow-up, adding to the clinician burden. Locally deployed language models could assist with routine chart review, but probability-based confidence fails to detect clinically important reasoning errors. We hypothesized that internal representational stability, how consistently a model’s hidden states respond to small input perturbations, provides a more reliable safety signal, enabling selective automation where uncertain cases are deferred to clinicians.
Materials/Methods: We developed Calibrating Confidence via Perturbation Stability (CCPS), which evaluates the stability of a model’s final hidden-layer representation across targeted perturbations of the same clinical note and question pair. Stable cases are automated; unstable cases are deferred for human review. Thresholds were tuned on a held-out training split (n=1,659) from an electronic health record QA corpus. CCPS was evaluated on an independent, multi-institutional test set of expert-labeled breast and pancreatic cancer notes (n=2,266) across three complexity tiers: Contextual (n=1,126), Synthesis (n=763), and Clinical Inference (n=377). Five open-weight, locally deployable models (0.5B to 3B parameters) were compared against probability-thresholding (PT) and internal-knowledge probing (PIK) baselines. Safety was defined as no greater than 5% incorrect answers among cases that the model automated rather than deferred.
Results: CCPS automated a larger share of cases than either baseline while keeping errors among non-deferred cases within the 5% threshold (Table 1). The next-best method (PIK) achieved lower yields of 32%, 13%, and 44% under the same constraint; PT failed to automate any cases in two tiers. CCPS produced the highest safe yield for every individual model in the Clinical Inference tier, and its advantage over baselines widened with task complexity.
Conclusion: Representational-stability calibration provides a principled safety layer for deploying small, locally hosted language models in oncology documentation review. By deferring uncertain cases rather than forcing predictions, this approach maintains strict error limits while reducing clinician review volume, supporting prospective evaluation in radiation oncology workflows.
Table 1. CCPS case-level triage across five locally hostable models. Automated cases are answered without human review; error rate reflects the proportion answered incorrectly. Median values; safety threshold: =5% error.| Complexity Tier | Total Cases | Automated by CCPS | Deferred to Clinician | Error Among Automated |
| Contextual | 1,126 | 405 (36%) | 721 (64%) | 4% |
| Synthesis | 763 | 114 (15%) | 649 (85%) | 4% |
| Clinical Inference | 377 | 189 (50%) | 188 (50%) | 2% |