Main Session
Sep 29
PQA 05 - Physics

3049 - Leveraging Multimodal Large Language Model for Clinical Information Extraction from Breast Cancer Examination Reports through Prompt Engineering

12:30pm - 01:45pm ET
Poster Hall - Exhibit Hall A
Screen: 32
POSTER

Presenter(s)

Jie Lan, PhD - West China Hospital of Sichuan University, Chengdu, Sichuan

J. Lan1,2, D. Gao3, X. Wu1,2, X. Chen1, X. Sun1, L. LIU1,2, L. Jia4, W. Zhang3, and J. Jing1; 1Institute of Breast Health Medicine, State Key Laboratory of Biotherapy, West China Hospital, Sichuan University and Collaborative Innovation Center, Chengdu, Sichuan, China, 2Division of Head and Neck Tumor Multimodality Treatment, Cancer Center, Institute of Breast Health Medicine, West China Hospital, Sichuan University, Chengdu, Sichuan, China, 3Shanghai United Imaging Healthcare Co., Ltd., Shanghai, China, 4Shenzhen United Imaging Healthcare Co., Ltd., Shenzhen, Guangdong, China

Purpose/Objective(s):

Extracting core clinical information from a large volume of patient reports to support breast cancer radiotherapy decision is an important part of radiation oncologists' (ROs) clinical work, which, however, is highly repetitive. To address this, in this study, by employing prompt engineering to simulate the cognitive process of ROs when reviewing reports, we proposes an automated clinical information extraction strategy based on a multimodal large language model (LLM).

Materials/Methods:

The core of the proposed extraction strategy lies in leveraging a multimodal LLM and prompt engineering to simulate the clinical information acquisition logic of ROs. The multimodal LLM used is Qwen-VL-Max. The first step of this strategy is to automatically identify and classify these different reports based on the characteristics of each report type. Following accurate classification of reports, the clinical information required to be extracted and the corresponding extraction rules are defined for each report category. The accuracy of extracted information is achieved through the iteration and optimization of prompts. Finally, the outputs are structured and constrained using the JSON format to ensure standardization of the extracted data.

Results:

For breast cancer, a total of ten distinct reports are involved, which can be grouped into six main categories: preoperative biopsy reports, imaging examination reports, chemotherapy reports, surgical records, postoperative pathology reports, and lymph node pathology reports. 36 types of clinical information are required to be extracted from these reports. The average accuracy rate for the 36 extracted types of clinical information exceeds 90%. Most of the extraction errors or omissions involve information with insufficiently clear directional descriptions in the reports, such as descriptions of tumors and benign nodules in ultrasound or MR reports, descriptions of surgical procedures in operative reports, and descriptions of tumor size in pathology reports. The lymph node information has the lowest extraction accuracy of less than 70%, primarily because the required output includes the total number of lymph nodes submitted for examination and the number of positive nodes, which involves summing the quantities from multiple batches or even multiple reports. Qwen-VL-Max shows relatively low accuracy when performing such calculations.

Conclusion:

Using multimodal LLM to extract clinical information from breast cancer patient reports through prompt engineering is potentially feasible and demonstrates the potential to reduce ROs' clinical workload.