Diagnostic Accuracy and Alignment With Human Reader Responses Across Five Large Language Models on Complex Clinical Case Challenges
DOI:
https://doi.org/10.54103/2282-0930/32167Abstract
Introduction: Large language models (LLMs) match or exceed physician accuracy on benchmark diagnostic tasks. Whether higher accuracy reflects human-like diagnostic behavior, which bears on how safely clinicians can supervise these systems, is rarely assessed.
Objectives: To determine whether diagnostic accuracy and alignment with human reader responses move together across widely accessible LLMs under default, point-of-care conditions.
Methods: This cross-sectional evaluation submitted the 29 most recent New England Journal of Medicine (NEJM) case challenges to 5 publicly accessible LLMs (GPT-5, Claude, Google Gemini, DeepSeek, and Grok) through each model’s web interface (November 5-8, 2025) at default settings, one query per case, no system prompt. Each case offered a clinical vignette and 6 candidate diagnoses; published human reader response distributions (4386-17 714 per case; median, 7732) were the comparator. Diagnostic accuracy against the published gold-standard diagnosis, compared with the human reader majority by the exact McNemar test; agreement by Cohen κ; alignment by a human-error similarity index (the share of model errors on cases the reader majority also missed) and by the Jensen-Shannon divergence (JSD) between each model’s response and the reader distribution; and ordinary least squares regression of per-case JSD on normalized reader-disagreement entropy.
Results: Reader majority accuracy was 48.3% (95% CI, 31.4-65.6). All 5 reached or exceeded this value, from 51.7% for DeepSeek (95% CI, 34.4-68.6) to 82.8% for Grok (95% CI, 65.5-92.4); Grok (P = .02), GPT-5 (P = .004), and Gemini (P = .04) significantly exceeded the reader majority, whereas Claude and DeepSeek did not. Grok, the most accurate model, showed the weakest alignment (similarity index, 0.40; mean JSD, 0.50 bits), with 3 of 5 errors on cases readers answered correctly. GPT-5 showed the strongest alignment (index, 1.000), failing only on cases that also challenged readers. For 4 of 5 models, divergence rose with reader-disagreement entropy; DeepSeek was the exception.
Conclusions and Relevance: In this evaluation, diagnostic accuracy and alignment with human reader responses diverged across LLMs. Because a model can outperform clinicians in aggregate yet fail where oversight is least likely to detect it, evaluation of clinical AI should weigh error structure alongside accuracy.
Downloads
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Rino Bellocco, Luca Soraci, Lorenzo Lo Cicero, Salvatore Silipigni, Giuseppe Lanfranchi, Savino Sciascia, Wisit Cheungpasitporn , Andrea Corsonello, Domenico Santoro, Guido Gembillo

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
How to Cite
Published 2026-09-22


