Quality and clinical implications of AI chatbot responses to common patient questions about total knee arthroplasty: a multidimensional comparative study.

Knee · Oct 03 2026 · Recent

Ekici M, Yalvaç ES, Okur KT, Demir EB, Demir MA, Üner Ç, et al.

Department of Orthopedics and Traumatology, Yozgat City Hospital, Yozgat, Türkiye

Adult Reconstruction

SUMMARY — THE REDUCTIONNine AI chatbots answering common total knee arthroplasty patient questions differed significantly in quality, usability, and transparency, so their usefulness for preoperative counseling varies and should be verified by clinicians.
Abstract, as published

INTRODUCTION: Patients considering total knee arthroplasty (TKA) frequently seek online guidance on indications, risks, recovery, rehabilitation, and precautions. While authoritative sources exist, generative AI chatbots are increasingly used for rapid answers, yet their quality, usability, and transparency may vary.

METHODS: In this cross-sectional comparative study, five common TKA patient FAQs were selected using clinic FAQs, AAOS OrthoInfo, search trends, and Delphi panel validation. Nine chatbots were accessed on 21 March 2026 using free public interfaces, in incognito sessions with cleared cookies. Queries were zero-shot, using only each question text. Each chatbot's five answers were pooled into a single "5-answer packet." Seven orthopaedic specialists, blinded to model identity, scored each packet using QUEST, DISCERN, PEMAT-Understandability/Actionability (PEMAT-U/A), JAMA Benchmarks, Trust (1-5), and Global Quality Score (GQS, 1-5). Inter-rater reliability was assessed via ICC (two-way random effects, absolute agreement, average measures). Model differences were tested using one-way ANOVA with Tukey HSD, reporting η2 effect sizes (α = 0.05).

RESULTS: ICC-average ranged 0.765-0.941. All scales differed significantly by model (ANOVA: QUEST F = 5.931; DISCERN F = 4.182; PEMAT-U F = 6.394; PEMAT-A F = 4.948; JAMA F = 17.083; Trust F = 4.031; GQS F = 4.477; all p ≤ 0.001; η2 0.374-0.717). DeepSeek had the highest DISCERN mean (70.57 ± 2.64), ChatGPT-4o the highest PEMAT-U/A (86.62 ± 4.45; 83.38 ± 9.81), and Microsoft Copilot the highest JAMA mean (2.00 ± 0.00).

CONCLUSION: AI chatbot performance for TKA FAQs is domain-dependent across quality, usability, and transparency. These findings have direct implications for orthopedic practice, particularly in preoperative patient counseling and shared decision-making in TKA.

Featured in the 2026-10-11 issue.

← Assessing the impact of dialysis on short, intermediate, and …Higher femoral periprosthetic fracture rate with dual-mobilit… →

The Reduction is a free email digest of newly published orthopaedic literature — a handful of new papers in the subspecialties you choose, each summarized like this one. Subscribe free or browse the archive.