Main Session
Sep 30
QP 35 - Digital Allies: AI Tools Built for the Real Clinical Environment

1210 - Modernizing the MD Anderson Stiefel Oropharynx Cancer Cohort Database through Development and Validation of Automated Data Extraction Algorithms

08:30am - 08:35am ET
Room 160

Presenter(s)

Waree Rinsurongkawong, PhD Headshot
Waree Rinsurongkawong, PhD - MD Anderson Cancer Center, Houston, TX

W. Rinsurongkawong1, A. Sahli2, K. Vidad2, C. Dede3, M. Mahin3, O. Starostina4, R. Lewis5, P. Roy4, M. Knafl6, W. N. Song2, K. A. Hutcheson2, J. J. Lee7, and A. C. Moreno8; 1The Department of Quantitative Research Computing, The University of Texas MD Anderson Cancer Center, Houston, TX, 2The Department of Head and Neck Surgery, The University of Texas MD Anderson Cancer Center, Houston, TX, 3The Department of Radiation Oncology, The University of Texas MD Anderson Cancer Center, Houston, TX, 4The Department of Enterprise Data Engineering and Analytics, The University of Texas MD Anderson Cancer Center, Houston, TX, 5The Department of Information Technology, The University of Texas MD Anderson Cancer Center, Houston, TX, 6The Department of Genomic Medicine, The University of Texas MD Anderson Cancer Center, Houston, TX, 7The Department of Biostatistics, The University of Texas MD Anderson Cancer Center, Houston, TX, 8Department of Radiation Oncology, The University of Texas MD Anderson Cancer Center, Houston, TX

Purpose/Objective(s): Manual curation of comprehensive, longitudinal clinical datasets can be time-consuming with limited scalability. While adoption of electronic health record (EHR) systems has greatly facilitated the exponential generation of electronic data, automated data extraction can be challenging due to a lack of standards and/or conflicting sources of data entry. The objective of this study was to develop and validate robust, high-quality algorithms for extracting multimodal data on patients enrolled in a single institution, prospective oropharyngeal cancer (OPC) registry.

Materials/Methods: Sociodemographic, treatment, and clinical outcomes data from patients enrolled in the MDA OPC registry (Protocol PA14-0947) from March 2015 to December 2024 were manually extracted to serve as a reference dataset. An integration and analytics platform leveraging EHR data was used to develop preprocessing modules for data element identification, retrieval, and transformation followed by data quality (DQ) assessment through a data validation module. DQ metrics including accuracy, recall, precision, and F1 scores were calculated and iteratively improved by adapting data mapping files in the preprocessing modules.

Results: Data from 1,817 OPC patients was automatically extracted and analyzed. The algorithm achieved 100% accuracy for age, sex, and race. Ethnicity and smoking status were extracted with completion rates of 99.8% and 99.6%, respectively, and demonstrated 100% concordance with manual curation. Vital status was 98.6% complete and accurate, with missing data in 22 patients. Approximately 20% of patients underwent a surgical procedure, while most received radiotherapy (83%), including concurrent chemotherapy (65%). Regex-based classification achieved high accuracy, recall, and precision for most surgical variables. In contrast, systemic therapy classification showed modestly lower performance, with F1 scores ranging from 0.76 for induction therapy to 0.87 for concurrent therapy.

Conclusion: Automated data extraction algorithms are both feasible and advantageous for improving the efficiency of data abstraction, particularly for data elements that are well structured within the EHR. This study underscores the importance of integrating DQ modules into extraction pipelines to characterize algorithm limitations, inform the development of data standards in priority domains, and support iterative performance improvement.