Main Session
Sep
28
PQA 03 - Digital Health Innovation and Informatics, Patient Safety & Quality, and Radiation and Cancer Biology
2499 - Structured Data from Free-Text Clinical Notes for the Assessment of Quality of Care: An AI-Assisted Abstraction and Validation Framework for Radiation Oncology
Presenter(s)
Rishabh Kapoor, PhD, MS, BS - Virginia Commonwealth University, Richmond, VA
R. Kapoor1,2, H. Hanfi1,2, J. R. Palta1, A. Roesner2, P. Turner2, M. D. Kelly1, and W. F. Sleeman1,2; 1National Radiation Oncology Program, Veterans’ Healthcare Administration, Washington, DC, 2Intelligent Health Solutions, Richmond, VA
Purpose/Objective(s):
To demonstrate the implementation of an AI-enabled clinical data abstraction pipeline from unstructured free-text clinical notes to score evidence-based clinical quality measures (CQMs) across a large radiation oncology network.Materials/Methods:
Clinical notes that included initial consultation, simulation, treatment planning, on-treatment visit, end of treatment, and follow-up, and anonymized DICOM/DICOM-RT treatment plans for 515 patients treated with moderate hypofractionation (60–70 Gy in 20–30 fractions) from 95 treating physicians were ingested into a structured large language models (LLM) pipeline. Google's MedGemma-27b LLM was locally deployed using domain, disease site, and encounter-specific prompts to abstract 133 clinical data elements per patient across all notes, 63 of which were used directly for CQM scoring. Prompt engineering with structured output constraints and terminology normalization addressed text variability. Extracted data underwent multilayer validation via system-level guardrails, clinical business rules, and predefined decision logic to detect inconsistencies, missing values, and implausible outputs. AI hallucinations and abnormal findings were flagged for human review to ensure source documentation fidelity. Curated datasets were processed through automated binary decision trees to generate pass/fail CQM scores, presented to providers via a dashboard. Results: The AI-assisted abstraction framework substantially reduced manual review burden compared with fully manual chart abstraction. LLM-specific failure modes included hallucinations, mis-parsed structured fields, and abbreviation ambiguity; these were mitigated by iterative encounter-specific prompt engineering templates, rule-based validators, and prioritized human adjudication. The finalized pipeline integrates data ingestion, AI abstraction, validation, human adjudication, and automated measure scoring into a reproducible workflow. While most data elements required for CQM scoring were successfully captured, certain elements remained challenging due to variable abbreviations, inconsistent terminology, or documentation dispersed across multiple clinical notes. Overall, the veracity of abstracted data improved substantially after application of layered validation and human adjudication, achieving high concordance with manual review for key CQM elements. Conclusion: An AI-driven abstraction pipeline supports scalable, enterprise-wide quality surveillance using existing clinical records, but unstructured note abstraction alone cannot ensure complete CQM element capture. A combined approach incorporating structured data entry through disease site–specific templates alongside AI-based abstraction from unstructured notes provides a more reliable data capture workflow. Ultimately, this pipeline empowers end users with seamless dashboard access to AI-extracted datasets, and actionable CQM performance scores.