Main Session
Sep 28
PQA 03 - Digital Health Innovation and Informatics, Patient Safety & Quality, and Radiation and Cancer Biology

2633 - Semantic Chunking and Dimensionality Reduction for Predicting Pathologic Complete Response to Neoadjuvant Chemotherapy from Clinical Narratives in Breast Cancer

10:45am - 12:00pm ET
Poster Hall - Exhibit Hall A
Screen: 26
POSTER

Presenter(s)

Kevin Tsang, MS - Cedars-Sinai Medical Center, Los Angeles, CA

K. Tsang1, H. Liu2, M. Kanthamneni3, J. Berkowitz1, N. Tatonetti1, and T. Dou3; 1Department of Computational Biomedicine, Cedars-Sinai Medical Center, West Hollywood, CA, 2Department of Hematology and Cellular Therapy, Cedars-Sinai Medical Center, West Hollywood, CA, 3Department of Radiation Oncology, Cedars-Sinai Medical Center, West Hollywood, CA

Purpose/Objective(s): Predicting pathologic complete response (pCR) to neoadjuvant chemotherapy (NAC) in breast cancer can inform treatment planning. Pre-treatment pathology and radiology reports contain rich prognostic information but are unstructured and exceed the input limits of embedding models. We developed a natural language processing pipeline using semantic chunking and dimensionality reduction to predict pCR from long clinical narratives.

Materials/Methods: We retrospectively studied 208 breast cancer patients treated with NAC; 46 (22.1%) achieved pCR based on manual annotation of post-treatment TNM staging. Pre-treatment reports were concatenated per patient (mean 43,950 tokens). Embeddings were generated using Qwen3-Embedding-4B (8,192-token limit). Narratives were partitioned into 50-token-overlapping 1,024-token chunks (mean 46 chunks/patient), and patient-level vectors were created via mean pooling. For anchor-guided retrieval, embeddings of clinically meaningful query concepts (e.g., “high-grade invasive carcinoma” and “triple-negative breast cancer”) were used to select the top-K most semantically relevant chunks by cosine similarity. To reduce overfitting (2,560 features; 208 samples), principal component analysis (PCA) preceded logistic regression. Performance was assessed using 5-fold cross-validated AUROC. We compared four strategies: A) truncated document-level embedding, B) chunked embedding with mean pooling, C) anchor-guided chunk retrieval, and D) chunked embedding with PCA reduction.

Results: Truncated document-level embeddings achieved AUROC 0.670. Chunking with mean pooling improved performance to 0.707. Anchor-guided retrieval (K=50) yielded AUROC 0.710. Best performance was achieved with mean-pooled chunks plus PCA to 100 components (AUROC 0.729). PCA consistently improved performance, with optimal results between 50–100 components.

Conclusion: Semantic chunking and PCA enable prediction of pCR from long pre-treatment clinical narratives, overcoming token-length constraints and reducing overfitting in small cohorts. This approach demonstrates the feasibility of leveraging routine clinical text for treatment response prediction without manual feature engineering and warrants validation in larger, multi-institutional datasets.

Table 1. Comparison of Text Embedding Strategies for Pathologic Complete Response (pCR) Prediction

Strategy

Description

AUROC (5-fold CV)

A) Truncated Document-level Embedding

Single embedding from report truncated to 8,192 tokens

0.670

B) Chunked Embedding with Mean Pooling

1,024-token overlapping chunks; pooled across all chunks

0.707

C) Anchor-guided Chunk Retrieval

Top-K clinically guided chunks before pooling

0.710

D) Chunk embedding with PCA Reduction

Mean-pooled chunks with PCA (50–100 components)

0.729