AI Clinical Decision Making Compared with Fellowship-Trained Hand Surgeons: A Case-Based Study.

J Hand Surg Am · Aug 19 2026 · Recent

Donnelley CA, Lee CJ, Halim A, Noback PC, Luo X

Department of Orthopaedics and Rehabilitation, Yale School of Medicine, New Haven, CT

Hand & Upper Extremity

SUMMARY — THE REDUCTIONComparing four AI platforms to fellowship-trained hand surgeons on clinical vignettes, all matched surgeon consensus on easy cases, but only Claude and Gemini remained reasonably concordant (60%) on complex cases versus ChatGPT-4 and Open Evidence (20%).
Abstract, as published

PURPOSE: Large language model-based artificial intelligence (AI) platforms have attracted substantial interest as potential clinical adjuncts within surgical specialties, particularly as patients have begun using AI for clinical advice. This study compared the management concordance of four commercially available AI platforms to practicing fellowship-trained hand surgeons using a series of clinical case vignettes within the field of hand surgery.

METHODS: Nine hand surgery clinical vignettes with multiple-choice answer options were developed, including representative radiographs and clinical images. Four cases were classified as straightforward ("easy") and five as clinically complex ("hard"). The cases were then distributed anonymously to fellowship-trained hand surgeons from three institutions using an online survey. Four AI platforms were queried using a standardized prompt: ChatGPT-4 (OpenAI), Claude Opus 4.6 (Anthropic), Gemini (Google DeepMind), and Open Evidence. All AI platforms were accessed in March 2026. Each platform was queried once per case without regeneration to reflect real-world single-query usage. The primary outcome was concordance between each AI platform and the surgeon plurality response for each case.

RESULTS: Fifteen fellowship-trained hand surgeons completed the survey. Surgeon consensus was high on easy cases (73% to 100% plurality agreement). All four AI platforms achieved concordance with the surgeon plurality on all four easy cases (100%). Performance diverged on complex cases: Claude Opus 4.6 and Gemini each achieved concordance on 3 of 5 complex cases (60%), whereas ChatGPT-4 and Open Evidence each achieved concordance on 1 of 5 (20%). Overall concordance was 7 of 9 (78%) for Claude Opus 4.6 and Gemini and 5 of 9 (56%) for ChatGPT-4 and Open Evidence.

CONCLUSIONS: AI platforms were in agreement with surgeon responses in the majority of low-complexity cases, but had reduced agreement in clinically complex cases.

Featured in the 2026-08-21 issue.

← Extensor pollicis longus rupture after both nonoperatively- a…The Articular Dislocation Angle of the First Metatarsophalang… →

The Reduction is a free email digest of newly published orthopaedic literature — a handful of new papers in the subspecialties you choose, each summarized like this one. Subscribe free or browse the archive.