Presenter(s)
R. Boggula1,2, A. Simko3, G. Heilemann4, and J. W. Burmeister1,2; 1Wayne State University School of Medicine, Detroit, MI, 2Karmanos Cancer Institute, Detroit, MI, 3Umeå University, Umeå, Sweden, 4Department of Radiation Oncology, Medical University of Vienna, Vienna, Austria
Purpose/Objective(s): To test the feasibility of an autonomous VMAT inverse-planning workflow in which an Anthropic Claude LLM proposes iterative updates to optimization weights and dose–volume objectives, with the goal of producing clinically acceptable plans.
Materials/Methods: Two previously treated prostate + pelvic lymph node cases (45 Gy in 25 fractions) planned with dual-arc VMAT were selected for this feasibility study. Clinical reference plans were created in Varian Eclipse (v16). CT images and structure sets were exported (DICOM) and imported into the open-source PyDoseRT platform for GPU-accelerated plan optimization. For each case, Anthropic Claude received the prescription and planning goals. Claude proposed initial optimization priorities (“weights”) for targets and organs-at-risk, and PyDoseRT then executed an optimization cycle (~2 minutes on an NVIDIA A100 GPU). After each cycle, Claude reviewed a standardized plan report (dose–volume metrics, goal pass/fail, and changes across iterations), reasoned the tradeoffs, and updated the optimization weights for the next cycle.
Results: The Claude-driven optimization loop produced clinically acceptable plans for both cases within 8–12 cycles. Metrics are reported as Claude-optimized vs Eclipse reference. In Case 1, Claude achieved target coverage meeting the clinical goal (PTV coverage = 95%), with substantially lower OAR dose (Bladder V40.5 Gy 13.8% vs 31.6%; Rectum V40.5 Gy 14.3% vs 25.1%; Sigmoid Dmax 46.2 Gy vs 47.7 Gy). In Case 2, Claude again met the target coverage goal (= 95%), with mixed OAR changes (Bladder V40.5 Gy 13.6% vs 10.2%; Rectum V40.5 Gy 8.4% vs 13.5%; Sigmoid Dmax 46.2 Gy vs 46.5 Gy). Review of the Claude logs showed adaptive behaviors consistent with planner reasoning (e.g., de-escalating target weights to resolve Dmax tradeoffs and accepting minor coverage changes to improve OAR sparing). Overall, Claude produced plan quality that was clinically comparable or superior to the Eclipse reference across both cases.
Conclusion: This proof-of-concept study shows that Anthropic Claude autonomously tuned inverse-planning parameters to generate VMAT plans that are comparable to or superior to clinical TPS reference plans. Larger multi-patient validation is needed to establish robustness and define the required clinical oversight.