Assessing Large Language Models for Clinical Coding in Hand Surgery: Effect of Note Authorship, Prompt Design, and Diagnosis/Procedure Type.

J Am Acad Orthop Surg · May 22 2026 · Recent

Schroeder AM, Goldenberg CB, Khaleel MI, Nuelle JAV, Kirby BJ, London DA

From the School of Medicine, University of Missouri-Columbia (Schroeder, Goldenberg, MO

Hand & Upper Extremity

SUMMARY — THE REDUCTIONLarge language models poorly predict ICD-10 codes (24% accuracy) but perform well for CPT codes (92%), requiring optimization before clinical use.
Abstract, as published

BACKGROUND: This study sought to assess large language models' (LLM) ability to generate correct ICD-10 and CPT codes using clinical documentation, and to determine whether note authorship, prompt design, or diagnosis/procedure type affect LLM performance. We hypothesized that LLMs can code hand surgery clinic and surgical notes with greater than 80% accuracy.

METHODS: Ninety patients evenly distributed across three orthopaedic hand surgeons and procedure types (cubital tunnel, carpal tunnel, and trigger finger release) were identified. Clinic and surgical notes were deidentified, and correct ICD-10 diagnosis and CPT procedure codes were recorded. "Zero-shot," "one-shot," "multishot," and "chain-of-thought" prompts instructed LLMs to assign ICD-10 codes and CPT codes based on note content. Each prompt was posed to Chat GPT 3.5, Chat GPT 4.0, and Gemini. Rates of coding correctness were calculated across attendings, diagnosis/procedure, prompt type, and LLM. Chi-square analysis determined statistical significance for these comparisons (P < 0.05).

RESULTS: No differences in LLM coding performance were observed between note authors (P = 0.09 ICD-10, P = 0.48 CPT) or prompt types (P = 0.27 ICD-10, P = 0.62 CPT). Chat GPT 3.5 provided less accurate ICD-10 codes than Chat GPT 4.0 or Gemini (P < 0.0001). All LLMs better predicted CPT codes (91.5% correct) than ICD-10 codes (23.9% correct). The most common error was incorrect or omitted ICD-10 laterality. Prompts updated to emphasize ICD-10 laterality demonstrated improved accuracy (40%).

DISCUSSION: Variation in note content and writing style did not markedly affect LLM performance. Public-facing LLMs require additional optimization to interpret clinical documentation for coding purposes and are not ready for independent use.

Featured in the 2026-06-26 issue.

← Strategies for Work-up and Treatment of Case Scenarios in Neu…Flexor Tendon Rupture after One Corticosteroid Injection in t… →

The Reduction is a free email digest of newly published orthopaedic literature — a handful of new papers in the subspecialties you choose, each summarized like this one. Subscribe free or browse the archive.