2542 - VOICE-AE - Voice-Based Oncology Intelligent Clinical Evaluation of Adverse Events: Using Ambient Audio and Large Language Models for Adverse Event Scoring from Clinical Encounters
Presenter(s)
F. Mastroleo1,2, M. Borras-Osorio2, J. Moonen2, J. A. Jordan2, K. K. Lin2, Q. Liu2, M. Zhou2, S. Shiraishi2, A. Y. K. Foong2, D. M. Routman2, and M. R. Waddle2; 1Department of Oncology and Hemato-Oncology, University of Milan, Milan, Italy, 2Department of Radiation Oncology, Mayo Clinic, Rochester, MN
Purpose/Objective(s): Treatment-related adverse events (AEs) are critical for assessing safety profiles, with Common Terminology Criteria for Adverse Events (CTCAE) as the defined national standard. Ambient audio use is increasing in clinical encounters and represents an untapped source of rich AE information. This pilot developed and evaluated a pipeline integrating speech recognition and Large Language Models (LLMs) to extract graded CTCAEs from clinician–patient audio recordings.
Materials/Methods: Clinical encounters of patients treated for prostate cancer were audio recorded and collected in a HIPAA-compliant environment, transcribed using OpenAI Whisper Large-v3 with WhisperX diarization, and processed through a multi-stage LLM pipeline: Gemini 2.5 Flash-lite for context-aware speaker recognition, Q&A extraction for AE identification, and Gemini 2.5 Pro for toxicity extraction mapping to CTCAE v5.0 grades (0–5). Six blinded raters of varying expertise independently graded 10 a priori-defined genitourinary/gastrointestinal CTCAE items. Hierarchical clustering (average linkage, distance = 1 - mean pairwise kappa) identified rater subgroups informing adjudication panel selection. Discrepancies underwent structured adjudication establishing the ground truth.
Results: Across the 39 patients (390 assessments), multi-rater Krippendorff's alpha among 6 human raters was 0.675. Pairwise weighted kappa averaged 0.586 (range: 0.476–0.709), with the highest agreement among the two radiation oncologists (RadOnc) and the Advanced Practice Provider (APP). Hierarchical clustering separated raters into two distinct subgroups: an expert clinician cluster (2 RadOnc + APP) and a trainee cluster (2 residents + research fellow), consistent with the expected effect of clinical expertise on CTCAE grading . The expert cluster served as the adjudication panel: 299/390 (76.7%) assessments were auto-agreed; 91 (23.3%) underwent structured REDCap adjudication. Against the adjudicated ground truth, Gemini 2.5 Pro achieved a mean weighted kappa of 0.717 (agreement 82.6%). Mean exact CTCAE grade match was 82.6% (range: 71.8%–100.0%), with 99.2% of predictions within one grade. Per-grade weighted-average F1 was 0.83. Hierarchical clustering positioned Gemini 2.5 Pro within the expert clinician cluster, indicating an AI grading pattern most closely resembling the most experienced raters.
Conclusion: Automated LLM-based CTCAE extraction from ambient audio achieved accuracy comparable to expert clinicians. This approach may enable automated documentation from routine encounters, reducing documentation burden while improving AE documentation completeness and standardization.