Main Session
Sep
28
PQA 03 - Digital Health Innovation and Informatics, Patient Safety & Quality, and Radiation and Cancer Biology
2665 - Automated Quality Auditing of Radiation Oncology Clinical Documentation Using Large Language Models
Presenter(s)
Bingyang Ye, PhD - Brigham and Women's Hospital, Cambridge, MA
B. Ye1, T. K. Kosak1, V. Goddla2, A. Murray1, M. Kearney1, N. E. Martin1, and D. S. Bitterman1; 1Brigham and Women's Hospital, Boston, MA, 2Harvard College, Cambridge, MA
Purpose/Objective(s):
Quality assurance (QA) of radiotherapy (RT) documentation is essential for patient safety and for meeting professional standards and accreditation requirements, but it requires substantial effort. For example, the American College of Radiology (ACR) Radiation Oncology Practice Accreditation requires evidence of chart documentation completeness, which currently requires extensive manual review because many required elements are present only in unstructured clinic notes. We developed and validated a large language model (LLM)-based system to automatically label each required documentation element as present or missing, and to classify each note as ‘complete’ versus ‘incomplete’.Materials/Methods:
We conducted a retrospective analysis of clinical notes from 25 patients treated with RT at one academic institution. The dataset included 210 notes (42 consult, 137 on-treatment visit (OTV), 31 completion notes) across 10 cancer types. We developed an LLM-based auditing system using GPT-4.1 mini to extract and evaluate 31 documentation elements across note types required for ACR accreditation (e.g., staging, pain assessment, on-treatment management, etc.). For note-level auditing, a note was labeled ‘complete’ if all required elements were present and ‘incomplete’ if one or more required element was missing. Departmental QA team members established ground truth and validated LLM-identified deficiencies. Manual chart review time was recorded for 5 patients.Results:
For element-level detection of missing documentation, the system achieved 94.2% accuracy, 59.3% precision, 86.9% recall, and 70.5% F1 across 1,058 element predictions. For note-level pass/fail classification, the system achieved 79.0% precision, 90.7% recall, and 84.5% F1. Manual chart review averaged 26 min/patient, totaling 10.8 hours for 25 patients. The LLM system processed all 210 notes in 58 minutes for an average 2.3 min/patient.Conclusion:
LLM-based documentation auditing can rapidly screen RT notes to identify charts most likely to be deficient, allowing clinicians to focus manual review where it is needed. The system caught most incomplete notes (91% recall) and kept false alarms limited (~1 in 5 flagged notes), while running ~11× faster than manual review. Performance was strongest for structured completion notes and lower for less standardized OTV documentation, supporting scalable QA workflows.| Note Type | Precision | Recall | F1 |
| Consult | 72.7% | 76.2% | 74.4% |
| OTV | 55.6% | 100% | 71.4% |
| Completion | 90.3% | 100% | 94.9% |