Diagnostic Accuracy and Alignment With Human Reader Responses Across Five Large Language Models on Complex Clinical Case Challenges

Authors

DOI:

https://doi.org/10.54103/2282-0930/32167

Abstract

Introduction: Large language models (LLMs) match or exceed physician accuracy on benchmark diagnostic tasks. Whether higher accuracy reflects human-like diagnostic behavior, which bears on how safely clinicians can supervise these systems, is rarely assessed.

Objectives: To determine whether diagnostic accuracy and alignment with human reader responses move together across widely accessible LLMs under default, point-of-care conditions.

Methods: This cross-sectional evaluation submitted the 29 most recent New England Journal of Medicine (NEJM) case challenges to 5 publicly accessible LLMs (GPT-5, Claude, Google Gemini, DeepSeek, and Grok) through each model’s web interface (November 5-8, 2025) at default settings, one query per case, no system prompt. Each case offered a clinical vignette and 6 candidate diagnoses; published human reader response distributions (4386-17 714 per case; median, 7732) were the comparator. Diagnostic accuracy against the published gold-standard diagnosis, compared with the human reader majority by the exact McNemar test; agreement by Cohen κ; alignment by a human-error similarity index (the share of model errors on cases the reader majority also missed) and by the Jensen-Shannon divergence (JSD) between each model’s response and the reader distribution; and ordinary least squares regression of per-case JSD on normalized reader-disagreement entropy.

Results: Reader majority accuracy was 48.3% (95% CI, 31.4-65.6). All 5 reached or exceeded this value, from 51.7% for DeepSeek (95% CI, 34.4-68.6) to 82.8% for Grok (95% CI, 65.5-92.4); Grok (P = .02), GPT-5 (P = .004), and Gemini (P = .04) significantly exceeded the reader majority, whereas Claude and DeepSeek did not. Grok, the most accurate model, showed the weakest alignment (similarity index, 0.40; mean JSD, 0.50 bits), with 3 of 5 errors on cases readers answered correctly. GPT-5 showed the strongest alignment (index, 1.000), failing only on cases that also challenged readers. For 4 of 5 models, divergence rose with reader-disagreement entropy; DeepSeek was the exception.

Conclusions and Relevance: In this evaluation, diagnostic accuracy and alignment with human reader responses diverged across LLMs. Because a model can outperform clinicians in aggregate yet fail where oversight is least likely to detect it, evaluation of clinical AI should weigh error structure alongside accuracy.

Downloads

Download data is not yet available.

Author Biographies

  • Rino Bellocco, University of Milano-Bicocca

    Department of Statistics and Quantitative Methods

  • Luca Soraci, Italian National Research Center on Aging (IRCCS INRCA), Cosenza,

    Unit of Geriatric Medicine and Clinical Epidemiology

  • Lorenzo Lo Cicero, University of Messina

    Unit of Nephrology, Department of Clinical and Experimental Medicine

  • Salvatore Silipigni, University of Messina

    Radiology Unit, Biomorf Department

  • Giuseppe Lanfranchi, University of Messina

    MIFT Department

  • Savino Sciascia, Ospedale San Giovanni Bosco

    University Center of Excellence on Nephrologic, Rheumatologic and Rare Diseases, ASL Città di Torino

  • Wisit Cheungpasitporn , Mayo Clinic

    Division of Nephrology and Hypertension, Mayo Clinic, Rochester

  • Andrea Corsonello

    Unit of Geriatric Medicine and Clinical Epidemiology

  • Domenico Santoro, University of Messina

    Unit of Nephrology, Department of Clinical and Experimental Medicine

  • Guido Gembillo, University of Messina

    Unit of Nephrology, Department of Clinical and Experimental Medicine

Downloads

Published

2026-09-22

How to Cite

1.
Diagnostic Accuracy and Alignment With Human Reader Responses Across Five Large Language Models on Complex Clinical Case Challenges. ebph [Internet]. 2026 Sep. 22 [cited 2026 Sep. 25]; Available from: https://riviste.unimi.it/index.php/ebph/article/view/32167
Received 2026-06-29
Published 2026-09-22