2675 - Development and Evaluation of a Specialized Large Language Model for Radiotherapy Knowledge Q&A
Presenter(s)
T. Zhang1, and L. Zhao2; 1Department of Radiation Oncology, Xijing Hospital, Fourth Military Medical University, Xi'an, Shaanxi, China, 2Department of Radiation Oncology, Xijing Hospital, Fourth Military Medical University, Xi’an, China
Purpose/Objective(s): General-purpose large language models (LLMs) often lack current, domain-specific knowledge and may generate unreliable content in specialized fields like radiation oncology. This study developed a specialized Q&A system to provide accurate, source-grounded knowledge support for radiotherapy.
Materials/Methods: A specialized knowledge base was built using authoritative resources, including AAPM reports, clinical guidelines (e.g., NCCN, CACA), textbooks, and literature. Texts underwent cleaning, deduplication, and semantic segmentation before being vectorized and indexed. The system implements an RAG architecture: for each query, relevant document snippets are retrieved and combined with the question to form an augmented prompt for the LLM. We compared a general model (DeepSeek as baseline) with a domain-adapted version fine-tuned on a curated radiotherapy Q&A dataset. Evaluation was performed by five senior radiation oncologists and medical physicists using 100 preset questions covering basic theory, clinical practice, and technical standards. Answers were scored on a 1–5 scale for accuracy and reliability, with ROUGE scores calculated as an auxiliary metric.
Results: The specialized system achieved an average score of 4.5 ± 0.7, significantly outperforming the general DeepSeek model (3.6 ± 1.0, p < 0.01). For questions involving specific quality control standards and technical parameters, the specialized system’s average score reached 4.7, with 95% of answers citing authoritative sources. The ROUGE-L score for the specialized system was 0.301, higher than the general model's 0.241.
Conclusion: Integrating a high-quality radiotherapy knowledge base with RAG and domain-adaptive fine-tuning significantly improves the accuracy, reliability, and source citation of LLMs in this specialized field. The developed system effectively reduces hallucinations and provides well-supported answers, demonstrating strong potential for clinical reference and decision support in radiotherapy.