2543 - Large Language Model-Generated Radiation Oncology Management Notes from Ambient Audio: A Blinded Comparative Evaluation
Presenter(s)
F. Mastroleo1,2, M. Borras-Osorio2, B. Kirchner2, S. Dobbs2, S. Shiraishi2, A. Y. K. Foong2, D. M. Routman2, and M. R. Waddle2; 1Department of Oncology and Hemato-Oncology, University of Milan, Milan, Italy, 2Department of Radiation Oncology, Mayo Clinic, Rochester, MN
Purpose/Objective(s): Weekly management notes impose significant documentation burden in radiation oncology. We evaluated whether large language models (LLMs) can generate clinically acceptable management notes directly from ambient audio of patient encounters, comparing two LLM variants against human-written notes.
Materials/Methods: Ambient audio transcriptions from prostate radiation weekly management visits were collected in a HIPAA-compliant setting and transcribed using automated speech recognition (Whisper Large-v3) with speaker diarization. Two LLMs, Gemini 2.5 Flash and Gemini 2.5 Pro, were deployed via Vertex AI (temperature=0.2) with structured JSON output to generate notes from clinically engineered prompts. Prompts incorporated the transcript, temporally filtered prior radiation oncology documentation, and treatment plan data from the electronic health record to prevent data leakage. Two blinded raters reviewed transcripts and independently scored all notes on a validated instrument across Completeness (1–5 Likert), Accuracy (1–5 Likert), Conciseness (1–5 Likert), Rewrite Needed (yes/no), and Rewrite Time (<1 min, 1–2 min, >2 min). Inter-rater scores were averaged. Paired Wilcoxon signed-rank (ordinal) and McNemar’s (binary) tests with Bonferroni correction across 10 comparisons assessed differences versus human notes.
Results: Across 32 encounters (96 total notes evaluated), inter-rater agreement was high (mean absolute difference =0.9 points across all dimensions). Both LLMs significantly outperformed human notes in Completeness (Flash: median 5.0, mean 4.88±0.28; Pro: 5.0, 4.70±0.40; human: 4.25, 4.17±0.70; p_corr=0.001 and 0.006, respectively). Pro achieved significantly higher Accuracy than human notes (median 5.0 vs 4.5, mean 4.94±0.17 vs 4.48±0.62, p_corr=0.002); Flash trended higher without reaching significance after correction (p_corr=0.32). Conciseness was comparable across all notes (medians 4.5–5.0). Human notes required rewriting in 75% of cases versus 25% for Pro (p_corr=0.014) and 41% for Flash (p_corr=0.24). Among notes requiring rewriting, LLM notes needed less editing: 85% of Flash rewrites (n=13) and 75% of Pro rewrites (n=8) completed in under one minute compared to 50% for human notes (n=24).
Conclusion: LLM-generated management notes from ambient audio demonstrated significantly greater completeness and reduced rewriting burden compared to human-written documentation, with Gemini 2.5 Pro performing best. LLM-assisted ambient audio note generation is promising for reducing documentation burden in radiation oncology while maintaining clinical quality.