2593 - Automated Clinical Phenotyping from Oncology Notes Using Local Retrieval-Augmented Language Model
Presenter(s)
P. Salome1,2, M. Knoll3, D. Walz4, N. Cogno1, A. S. Dedeoglu5, A. L. Qi1, S. J. Isakoff6, A. Abdollahi7, R. B. Jimenez8, D. S. Bitterman9, H. Paganetti1, and I. Chamseddine1; 1Department of Radiation Oncology, Massachusetts General Hospital/Mass General Brigham and Harvard Medical School, Boston, MA, 2Clinical Cooperation Unit Translational Radiation Oncology, German Cancer Research Center (DKFZ), Heidelberg, Germany, 3Division of Molecular and Translational Radiation Oncology, Department of Radiation Oncology, Heidelberg Faculty of Medicine (MFHD), Heidelberg University Hospital (UKHD) and Heidelberg Ion-Beam Therapy Center (HIT), Heidelberg, Germany, 4Faculty of Medicine (MFHD), Heidelberg University Hospital (UKHD),Division of Molecular and Translational Radiation, German Cancer Consortium (DKTK) Core-Center Heidelberg, National Center for Tumor Diseases (NCT), Heidelberg Ion-Beam Therapy Center (HIT), Heidelberg, Germany, 5MGH Cancer Center, Boston, MA, 6Massachusetts General Hospital, Boston, MA, 7Clinical Cooperation Unit Translational Radiation Oncology, National Center for Tumor Diseases (NCT), Heidelberg University Hospital (UKHD) and German Cancer Research Center (DKFZ), Heidelberg, Germany, 8Department of Radiation Oncology, Massachusetts General Hospital, Boston, MA, 9Department of Radiation Oncology, Mass General Brigham/Dana-Farber Cancer Institute, Harvard Medical School, Boston, MA
Purpose/Objective(s): Manual extraction of clinical variables from oncology notes is a critical bottleneck for registries, quality reporting, and clinical research. Existing automated approaches rely on large cloud-hosted proprietary models, necessitating external data transfer, or require task-specific fine-tuning on curated datasets. We introduce a graph-based retrieval-augmented generation pipeline paired with a locally deployed mid-size language model with the objective to achieve extraction accuracy comparable to manual curation without fine-tuning or external data sharing.
Materials/Methods: We developed a four-phase pipeline that (1) generates feature-specific search terms via ontology enrichment, (2) constructs a clinical knowledge graph using biomedical named entity recognition, (3) retrieves context through graph-diffusion reranking, and (4) extracts features via structured prompts. The pipeline ran locally using a 14B-parameter open-source model requiring no external data transfer. It was applied to three cohorts: triple-negative breast cancer (TNBC; n=104, 42 features; primary development), recurrent high-grade glioma (n=191, 19 features; cross-lingual validation in German language), and a public critical care dataset (n=100, 10 features; zero-shot external validation). Performance was assessed by macro F1 scores against manual annotation. Clinical utility was evaluated for the TNBC cohort by comparing Cox models for 3-year progression-free survival built from automated versus manual features (Harrell C-index).
Results: F1 scores were 0.80±0.11 (TNBC; 44 patients, 42 features), 0.79±0.13 (glioma; 61 patients, 19 features), and 0.84±0.07 (external validation; 100 patients, 10 features). Compared to direct prompting and naive RAG baselines, F1 improved by 0.14–0.21 and 0.16–0.17. Manual configuration refinement further improved F1 to 0.83 (TNBC) and 0.81 (glioma). Extraction averaged 1.7–1.9 seconds per feature; complete TNBC extraction (42 features, 104 patients) took under 2.5 hours versus approximately two weeks manually. A smaller 3.8B model reduced time by 57% with an F1 decrease of 0.03–0.10. Survival models from automated features performed comparably to manual curation (C-index 0.77 vs. 0.76; p=0.512).
Conclusion: This locally deployed pipeline achieves accurate automated feature extraction from multilingual oncology notes without fine-tuning or external data sharing. Validated across three cohorts and two languages, it produced survival models indistinguishable from manual curation. By reducing abstraction from weeks to hours while preserving data sovereignty, this approach can accelerate radiation oncology outcomes research, registry participation, and quality improvement.