Editor’s summary
Computational tools for medical decision support have been advancing over time, mainly by serving as resources for limited applications. Machine learning tools for autonomous interpretation of clinical cases have also been gradually improving over time. Brodeur et al. pitted a large language model, the OpenAI o1 series, directly against hundreds of physicians at different levels of training and experience on a variety of clinical cases ranging from published patient vignettes to evaluations of brand-new emergency room patients, as well as on clinical tasks including both diagnosis and planning of clinical management (see the Perspective by Hopkins and Cornelisse). Across a variety of scenarios and applications, the large language model outperformed both human physicians and older models, suggesting its potential utility for clinical care. —Yevgeniya Nusinovich
Abstract
More than 65 years ago, complex clinical diagnostic reasoning cases were introduced as the gold standard for the evaluation of expert medical computing systems, a standard that has held ever since. In this study, we report the results of a physician evaluation of a large language model (LLM) on challenging clinical cases across five experiments with a baseline of hundreds of physicians. We then report a real-world study comparing human expert and artificial intelligence (AI) second opinions in randomly selected patients in the emergency room of a major tertiary academic medical center. In all experiments, the LLM outperformed physician baselines and displayed continued improvement from prior generations of AI clinical decision support. Our study suggests that LLMs have eclipsed most benchmarks of clinical reasoning, motivating the urgent need for prospective trials.
Access the full article
View all access options to continue reading this article.
Supplementary Materials
The PDF file includes:
Other Supplementary Material for this manuscript includes the following:
MDAR Reproducibility Checklist
- Download
- 287.16 KB
References and Notes
1
R. S. Ledley, L. B. Lusted, Reasoning foundations of medical diagnosis; symbolic logic, probability, and value theory aid our understanding of how physicians reason. Science 130, 9–21 (1959).
2
K. Brodman, A. J. Erdmann Jr., I. Lorge, C. P. Gershenson, H. G. Wolff, The Cornell Medical Index-Health Questionnaire. III. The evaluation of emotional disturbances. J. Clin. Psychol. 8, 119–124 (1952).
3
F. T. de Dombal, D. J. Leaper, J. R. Staniland, A. P. McCann, J. C. Horrocks, Computer-aided diagnosis of acute abdominal pain. BMJ 2, 9–13 (1972).
4
E. Shortliffe, “Mycin: A knowledge-based computer program applied to infectious diseases”in Proceedings: Symposium on Computer Applications in Medical Care, Washington, DC, 3 to 5 October 1977 (Hanley & Belfus, 1977), pp. 66–69.
5
E. B. Ing, M. Balas, G. Nassrallah, D. DeAngelis, N. Nijhawan, The Isabel differential diagnosis generator for orbital diagnosis. Ophthalmic Plast. Reconstr. Surg. 39, 461–464 (2023).
6
R. A. Miller, H. E. Pople Jr., J. D. Myers, Internist-1, an experimental computer-based diagnostic consultant for general internal medicine. N. Engl. J. Med. 307, 468–476 (1982).
7
Z. Kanjee, B. Crowe, A. Rodman, Accuracy of a Generative Artificial Intelligence Model in a Complex Diagnostic Challenge. JAMA 330, 78–80 (2023).
8
D. McDuff, M. Schaekermann, T. Tu, A. Palepu, A. Wang, J. Garrison, K. Singhal, Y. Sharma, S. Azizi, K. Kulkarni, L. Hou, Y. Cheng, Y. Liu, S. S. Mahdavi, S. Prakash, A. Pathak, C. Semturs, S. Patel, D. R. Webster, E. Dominowska, J. Gottweis, J. Barral, K. Chou, G. S. Corrado, Y. Matias, J. Sunshine, A. Karthikesalingam, V. Natarajan, Towards accurate differential diagnosis with large language models. Nature 642, 451–457 (2025).
9
H. Nori, N. Usuyama, N. King, S. M. McKinney, X. Fernandes, S. Zhang, E. Horvitz, From Medprompt to o1: Exploration of run-time strategies for medical challenge problems and beyond. arXiv:2411.03590 [cs.CL] (2024).
11
P. Lee, S. Bubeck, J. Petro, Benefits, limits, and risks of GPT-4 as an AI chatbot for medicine. N. Engl. J. Med. 388, 1233–1239 (2023).
12
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, H. Nori, H. Palangi, M. T. Ribeiro, Y. Zhang, Sparks of Artificial General Intelligence: Early experiments with GPT-4. arXiv:2303.12712 [cs.CL] (2023).
13
S. Johri, J. Jeong, B. A. Tran, D. I. Schlessinger, S. Wongvibulsin, L. A. Barnes, H.-Y. Zhou, Z. R. Cai, E. M. Van Allen, D. Kim, R. Daneshjou, P. Rajpurkar, An evaluation framework for clinical use of large language models in patient interaction tasks. Nat. Med. 31, 77–86 (2025).
14
A. Rodman, L. Zwaan, A. Olson, A. K. Manrai, When it comes to benchmarks, humans are the only way. NEJM AI 2, AIe2500143 (2025).
15
R.-E. E. Abdulnour, A. S. Parsons, D. Muller, J. Drazen, E. J. Rubin, J. Rencic, Deliberate practice at the virtual bedside to improve clinical reasoning. N. Engl. J. Med. 386, 1946–1947 (2022).
16
S. Cabral, D. Restrepo, Z. Kanjee, P. Wilson, B. Crowe, R.-E. Abdulnour, A. Rodman, Clinical Reasoning of a Generative Artificial Intelligence Model Compared With Physicians. JAMA Intern. Med. 184, 581–583 (2024).
17
V. Schaye, L. Miller, D. Kudlowitz, J. Chun, J. Burk-Rafel, P. Cocks, B. Guzman, Y. Aphinyanaphongs, M. Marin, Development of a Clinical Reasoning Documentation Assessment Tool for Resident and Fellow Admission Notes: A Shared Mental Model for Feedback. J. Gen. Intern. Med. 37, 507–512 (2022).
18
E. Goh, R. J. Gallo, E. Strong, Y. Weng, H. Kerman, J. A. Freed, J. A. Cool, Z. Kanjee, K. P. Lane, A. S. Parsons, N. Ahuja, E. Horvitz, D. Yang, A. Milstein, A. P. J. Olson, J. Hom, J. H. Chen, A. Rodman, GPT-4 assistance for improvement of physician performance on patient care tasks: A randomized controlled trial. Nat. Med. 31, 1233–1238 (2025).
19
E. Goh, R. Gallo, J. Hom, E. Strong, Y. Weng, H. Kerman, J. A. Cool, Z. Kanjee, A. S. Parsons, N. Ahuja, E. Horvitz, D. Yang, A. Milstein, A. P. J. Olson, A. Rodman, J. H. Chen, Large language model influence on diagnostic reasoning: A randomized clinical trial: A randomized clinical trial. JAMA Netw. Open 7, e2440969 (2024).
20
E. S. Berner, G. D. Webster, A. A. Shugerman, J. R. Jackson, J. Algina, A. L. Baker, E. V. Ball, C. G. Cobbs, V. W. Dennis, E. P. Frenkel, L. D. Hudson, E. L. Mancall, C. E. Rackley, O. D. Taunton, Performance of four computer-based diagnostic systems. N. Engl. J. Med. 330, 1792–1796 (1994).
21
D. J. Morgan, L. Pineles, J. Owczarzak, L. Magder, L. Scherer, J. P. Brown, C. Pfeiffer, C. Terndrup, L. Leykum, D. Feldstein, A. Foy, D. Stevens, C. Koch, M. Masnick, S. Weisenberg, D. Korenstein, Accuracy of practitioner estimates of probability of diagnosis before and after testing. JAMA Intern. Med. 181, 747–755 (2021).
22
R. M. Ratwani, D. W. Bates, D. C. Classen, Patient safety and artificial intelligence in clinical care. JAMA Health Forum 5, e235514 (2024).
23
Q. Jin, F. Chen, Y. Zhou, Z. Xu, J. M. Cheung, R. Chen, R. M. Summers, J. F. Rousseau, P. Ni, M. J. Landsman, S. L. Baxter, S. J. Al’Aref, Y. Li, A. Chen, J. A. Brejt, M. F. Chiang, Y. Peng, Z. Lu, Hidden flaws behind expert-level accuracy of multimodal GPT-4 vision in medicine. NPJ Digit. Med. 7, 190 (2024).
24
D. E. Newman-Toker, S. M. Peterson, S. Badihian, A. Hassoon, N. Nassery, D. Parizadeh, L. M. Wilson, Y. Jia, R. Omron, S. Tharmarajah, L. Guerin, P. B. Bastani, E. A. Fracica, S. Kotwal, K. A. Robinson, Agency for Healthcare Research and Quality, US Department of Health and Human Services, “Diagnostic errors in the emergency department: A systematic review” (AHRQ, 2022); https://effectivehealthcare.ahrq.gov/products/diagnostic-errors-emergency-updated/research.
25
A. D. Auerbach, T. M. Lee, C. C. Hubbard, S. R. Ranji, K. Raffel, G. Valdes, J. Boscardin, A. K. Dalal, A. Harris, E. Flynn, J. L. Schnipper, UPSIDE Research Group, Diagnostic errors in hospitalized adults who died or were transferred to intensive care. JAMA Intern. Med. 184, 164–173 (2024).
26
T. Buckley, J. A. Diao, P. Rajpurkar, A. Rodman, A. K. Manrai, Multimodal Foundation Models Exploit Text to Make Medical Image Predictions arXiv:2311.05591 [cs.CV] (2024).
27
T. A. Buckley, R. Conci, P. G. Brodeur, J. Gusdorf, S. Beltrán, B. Behrouzi, B. Crowe, J. Dockterman, M. Muhammad, S. Ohnigian, A. Sanchez, J. A. Diao, A. P. Shah, D. Restrepo, E. S. Rosenberg, A. S. Lea, M. Zitnik, S. H. Podolsky, Z. Kanjee, R.-E. E. Abdulnour, J. M. Koshy, A. Rodman, A. K. Manrai, Advancing medical artificial intelligence using a century of cases arXiv:2509.12194 [cs.AI] (2025).
28
F. Yu, A. Moehring, O. Banerjee, T. Salz, N. Agarwal, P. Rajpurkar, Heterogeneity and predictors of the effects of AI assistance on radiologists. Nat. Med. 30, 837–849 (2024).
29
G. Dhaliwal, C. M. Hood, A. K. Manrai, T. A. Buckley, A. W. Asombang, E. L. Hohmann, Case 28-2025: A 36-year-old man with abdominal pain, fever, and hypoxemia. N. Engl. J. Med. 393, 1421–1434 (2025).
30
M. Goldszmidt, J. P. Minda, G. Bordage, Developing a unified list of physicians’ reasoning tasks during clinical encounters. Acad. Med. 88, 390–397 (2013).
32
W. F. Bond, L. M. Schwartz, K. R. Weaver, D. Levick, M. Giuliano, M. L. Graber, Differential diagnosis generators: An evaluation of currently available computer programs. J. Gen. Intern. Med. 27, 213–219 (2012).
33
P. Fritz, A. Kleinhans, R. Raoufi, A. Sediqi, N. Schmid, S. Schricker, M. Schanz, C. Fritz-Kuisle, P. Dalquen, H. Firooz, G. Stauch, M. D. Alscher, Evaluation of medical decision support systems (DDX generators) using real medical cases of varying complexity and origin. BMC Med. Inform. Decis. Mak. 22, 254 (2022).
34
A. Rodman, T. A. Buckley, A. K. Manrai, D. J. Morgan, Artificial intelligence vs clinician performance in estimating probabilities of diagnoses before and after testing. JAMA Netw. Open 6, e2347075 (2023).
35
J. Gallifant, M. Afshar, S. Ameen, Y. Aphinyanaphongs, S. Chen, G. Cacciamani, D. Demner-Fushman, D. Dligach, R. Daneshjou, C. Fernandes, L. H. Hansen, A. Landman, L. Lehmann, L. G. McCoy, T. Miller, A. Moreno, N. Munch, D. Restrepo, G. Savova, R. Umeton, J. W. Gichoya, G. S. Collins, K. G. M. Moons, L. A. Celi, D. S. Bitterman, The TRIPOD-LLM reporting guideline for studies using large language models. Nat. Med. 31, 60–69 (2025).