Towards conversational diagnostic artificial intelligence
Abstract At the heart of medicine lies physician–patient dialogue, where skillful history-taking enables effective diagnosis, management and enduring trust 1,2 . Artificial intelligence (AI) systems capable of diagnostic dialogue could increase accessibility and quality of care. However, approximating clinicians’ expertise is an outstanding challenge. Here we introduce AMIE (Articulate Medical Intelligence Explorer), a large language model (LLM)-based AI system optimized for diagnostic dialogue. AMIE uses a self…
Hospital Universitari Doctor Peset, València 03 by 19Tarrestnom65. CC BY-SA 4.0 · https://creativecommons.org/licenses/by-sa/4.0
Researchers introduced AMIE, an LLM-based system for diagnostic dialogue, and tested it against 20 primary care physicians in 159 text-based scenarios with patient-actors from Canada, the UK and India. Specialist and patient-actor raters scored performance across 32 and 26 axes including history-taking and management.
Outperforming physicians in a simulated text setting suggests potential to increase accessibility and quality of care, but the unfamiliar chat modality and simulated environment leave open whether gains would persist in real-world clinical workflows, in-person care, or diverse patient populations.
- AMIE is a large language model-based system optimized for diagnostic dialogue using self-play-based simulated environment with automated feedback.
- Evaluation framework covered history-taking, diagnostic accuracy, management, communication skills and empathy.
- Study included 159 case scenarios from providers in Canada, the United Kingdom and India with 20 primary care physicians compared to AMIE.
In a randomized double-blind crossover study of text consultations, AMIE showed greater diagnostic accuracy than primary care physicians.
The rundown
The study design was a randomized, double-blind crossover of text-based consultations modeled on objective structured clinical examination with validated patient-actors. Evaluations were performed separately by specialist physicians and by patient-actors.
AMIE was trained to scale learning across disease conditions, specialties and contexts. Authors described results as a milestone towards conversational diagnostic AI while noting caution in interpretation.
Sources
- Peer-reviewedNature2025-04-09
How should this claim be treated?
ace
The debate