Presenter(s)
S. Cui1, B. Bioren2, J. Hellerstein3, F. Yaseen4, J. Kang1, Y. He1, S. R. Bowen1, and J. Fu1; 1Department of Radiation Oncology, University of Washington/Fred Hutchinson Cancer Center, Seattle, WA, 2School of Computer Science & Engineering, University of Washington, Seattle, WA, 3eScience Institute, University of Washington, Seattle, WA, 4Department of Biomedical Informatics and Medical Education, University of Washington, Seattle, WA
Purpose/Objective(s): Pathology reports contain rich, unstructured clinical information that may inform prognosis in non-small-cell lung cancer (NSCLC). Large language models (LLMs) provide an opportunity to extract predictive signals directly from free-text clinical reports. This study evaluated LLM-based approaches for predicting overall survival (OS) in NSCLC using pathology reports.
Materials/Methods: Pathology reports and follow-up data were obtained from the TCGA LUAD and LUSC cohorts. Of 1089 initially identified NSCLC patients, those with multiple reports, missing survival data, or inadequate follow-up time were excluded, resulting in a final cohort of 658 patients. Two-year survival was modeled as a binary outcome (=2-year survival vs. death within 2 years), a clinically meaningful landmark frequently reported in NSCLC studies. Patients were stratified by the survival status and split into training (70%), validation (15%), and independent test (15%) sets. Two modeling paradigms were evaluated: (1) an end-to-end zero-shot pre-trained LLM (Gemini v2.5 Flash) that generated survival probabilities directly from raw pathology text using a structured, task-specific prompt, and (2) an embedding-based framework in which Jina-V3 representations were extracted and used for downstream random forest classification. Within the embedding-based framework, both pre-trained and fine-tuned Jina-V3 models were assessed. Fine-tuning was performed using parameter-efficient low-rank adaptation (LoRA) and supervised online contrastive loss to optimize outcome-specific embedding separation. Model performance was evaluated on the independent test set using the area under the ROC curve (AUC), with pairwise comparisons conducted using DeLong’s test.
Results: The zero-shot Gemini model achieved an AUC of 0.641 (95% CI: 0.524-0.757) on the independent test set. The embedding-based pipeline using pre-trained Jina-V3 achieved an AUC of 0.615 (95% CI: 0.493-0.736). Fine-tuned Jina-V3 embeddings demonstrated the highest discrimination (AUC 0.675, 95% CI: 0.558-0.793). Although fine-tuning yielded numerically improved performance, differences between models were not statistically significant.
Conclusion: LLM-based approaches demonstrated moderate discrimination for predicting OS in NSCLC from pathology reports. Zero-shot generative LLMs achieved performance comparable to traditional embedding-based pipelines, supporting the feasibility of streamlined prognostic modeling directly from free text. Task-specific fine-tuning of embeddings using parameter-efficient contrastive learning provided incremental improvement, suggesting that adaptation enhances prognostic signal extraction. Future work will integrate multimodal features, including structured clinical variables and imaging data, to potentially enhance prognostic accuracy.