BACKGROUND: Patients use artificial intelligence-based large language models (AI-LLMs) to research minimally invasive surgery (MIS) for hallux valgus; however, their reliability remains uninvestigated.
PURPOSE: To compare the quality and readability of ChatGPT and Gemini responses regarding MIS bunion surgery.
METHODS: Ten frequently asked questions reflecting diverse clinical and procedural inquiries were submitted to ChatGPT and Gemini. Quality was assessed via DISCERN and 5-point Likert scales. Readability and actionability were evaluated using the Patient Education Materials Assessment Tool (PEMAT) and Flesch-Kincaid Reading Ease (FKRE).
RESULTS: No significant differences existed between models (p > 0.05). Both exceeded sufficiency thresholds for DISCERN (49.7; 49) and Likert (5.0; 4.9). While understandability (85.9%; 83.3%) and FKRE (32.9; 33.5) met requirements, both failed actionability (38.3%; 34%).
CONCLUSION: Both AI models offer reliable, high-quality theoretical information regarding MIS hallux valgus surgery. However, they are insufficient in providing actionable guidance and exceed ideal reading complexity for general patient populations.
Read the article: PubMed · Publisher (DOI)