Exploring large language models for specialist-level oncology care

Vikram Dhillon
Polly Niravath
Preethi Prasad
Khaled Saab
Ryutaro Tanno
Yong Cheng
Hanh Mai
Ethan Burns
Zainub Ajmal
Kavita Kulkarni
Philip Mansfield
Joelle Barral
Juro Gottweis
Sara Mahdavi
NEJM AI (2025)

Abstract

Large language models have shown rapid progress in encoding clinical knowledge and demonstrating clinical reasoning. However, their capabilities in subspecialty or complex medical settings remain underexplored. In this work, we probe the performance of Articulate Medical Intelligence Explorer (AMIE), a conversational diagnostic AI system in the subspecialty of breast oncology care without specific fine-tuning to this challenging domain. To perform this evaluation, we curated a set of 60 synthetic breast cancer vignettes representing a range of treatment-naive, treatment-refractory, and rare histology cases encountered in a community-based breast oncology clinic. We developed a detailed clinical rubric for evaluating management plans, including axes such as the quality of case summarization, safety of the proposed care plan, and recommendations for treatment (i.e., chemotherapy, radiotherapy, surgery, and hormonal therapy). To improve performance, we enhanced AMIE with the inference-time ability to perform web search retrieval to gather relevant and up-to-date clinical knowledge and refine its responses with a multistage, self-critique pipeline. We compare the response quality of AMIE with that of internal medicine trainees, oncology fellows, and general oncology attendings under both automated and specialist clinician evaluations. Although our evaluations were limited to a few physicians, AMIE outperformed trainees and fellows, demonstrating the potential of the system in this important domain. However, AMIE’s performance was overall inferior to that of attending oncologists, suggesting that further prospective research is needed.

Research Areas

×