288 - Evaluation of Large Language Model Performance in Extracting Structured Oncologic Variables from Head and Neck Cancer Pathology Reports
Presenter(s)
V. Chennupati1, M. Mallick1, D. Alvarez2, E. Gogineni3, S. Baliga4, D. J. Konieczkowski3, D. L. Mitchell3, S. R. Jhawar5, J. C. Grecula3, P. Bhateja6, L. Miller7, J. W. Rocco7, M. Bonomi6, D. M. Blakaj3, S. J. Ma2, and S. Zhu8; 1The Ohio State University College of Medicine, Columbus, OH, 2The Ohio State University, Columbus, OH, 3Department of Radiation Oncology, James Cancer Hospital/Wexner Medical Center, The Ohio State University, Columbus, OH, 4Department of Radiation Oncology, James Cancer Hospital, The Ohio State University Medical Center, Columbus, OH, 5Department of Radiation Oncology, The James Cancer Center, Ohio State University Wexner Medical Center, Columbus, OH, 6The Ohio State University Wexner Medical Center, Columbus, OH, 7The Ohio State University Department of Otolaryngology - Head & Neck Surgery, Columbus, OH, 8University of Florida, Gainesville, FL
Purpose/Objective(s):
Clinical research in postoperative head and neck cancer (HNC) frequently requires the creation of datasets derived from surgical pathology reports, but manual data extraction is time-consuming and prone to error. In this study, we evaluated the efficacy of a large language model (LLM) for automating this task.
Materials/Methods:
In an IRB-approved retrospective study, 251 patients who underwent surgical resection for HNC were included. Ground truth values for eleven pathologic variables were manually extracted from pathology reports by two medical students and reviewed by two board-certified radiation oncologists for quality assurance. Extracted variables included neck dissection laterality (left, right, bilateral, or none), tumor differentiation (well, moderately, or poorly differentiated, or undifferentiated), depth of invasion (cm), lymphovascular invasion (LVI), perineural invasion (PNI), extranodal extension (ENE), total lymph node count, positive lymph node count, final margin status (positive or negative), and pathologic AJCC T and N classifications (pT and pN). Continuous variables (depth of invasion, total lymph node count, and positive lymph node count) were recorded numerically, while categorical variables were coded as integers to facilitate downstream evaluation. An open-source LLM, Qwen3-8B, was prompted to extract these variables in the same format as the ground truth. Accuracy was calculated independently for each variable as the primary performance metric.
Results:
Across 251 cases, the LLM demonstrated high concordance with physician-validated manual extraction for most variables. Accuracy was 100% for total lymph node count, ENE, and LVI; 99.6% for pN; 98.8% for positive lymph node count; 96.4% for tumor differentiation; 95.2% for pT; 90.4% for depth of invasion; 85.7% for neck dissection laterality; and 81.7% for final margin status.
Conclusion:
The LLM achieved high accuracy in extracting structured HNC pathologic variables from surgical pathology reports, particularly for nodal staging, invasion variables, and lymph node counts, with comparatively lower performance for margin status and laterality of neck dissection. These findings support further investigation of LLM-augmented chart review as a tool to accelerate research data collection while maintaining accuracy.