Direct Comparison Between Loss-to-Follow-Up and Statistical Fragility Is Methodologically Inappropriate, and Fragility Reflects the P Value, Not Trial Robustness: A Simulation Analysis of 300,000 Randomized Controlled Trials.

Arthroscopy · Aug 16 2026 · Recent

Vivekanantha P, Son H, Kay J, Kotipalli S, Bouchard MD, Madden K, et al.

Division of Orthopaedic Surgery, Department of Surgery, McMaster University, Canada

General Orthopaedics

SUMMARY — THE REDUCTIONSimulation of 300,000 RCTs shows fragility indices are mathematically driven by p-values, not true trial robustness, and shouldn't be directly compared to loss-to-follow-up thresholds.
Abstract, as published

PURPOSE: To assess the relation between the fragility index (FI) and reverse fragility index (RFI) with the minimum number of patients needed to reverse statistical significance (e.g. henceforth termed the lost to follow-up index (LTFI) and reverse LTFI (R-LTFI), respectively) and apply machine learning to identify which trial parameters are most important in FI, RFI, and continuous fragility index composition given the nonlinearity of these metrics.

METHODS: A total of 300,000 randomized controlled trials (100,000 for each metric) were simulated using common trial parameter value ranges. For FI and RFI, LTFI and R-LTFI values, respectively, were calculated as the minimum number of patients lost to follow-up to reverse significance in either direction. Machine learning models were trained to assess the relative importance of P value, sample size, and event numbers in FI and RFI composition and P value, group means, group standard deviations, and group sizes for continuous fragility index composition.

RESULTS: Among their respective cohorts of 100,000, the LTFI and R-LTFI were greater than FI and RFI in 84.9% and 95.5% of simulated randomized controlled trials, respectively. Random Forest and XGBoost machine learning models had near perfect accuracy in modeling fragility metrics (R2 ≥ 0.98), justifying its use over standard linear regression models in identifying trial parameters that are most important in determining fragility values. Feature importance analysis showed that the P value accounted for 79.8%, 71.9%, and 64.5% of variability in FI, RFI, and continuous fragility index values.

CONCLUSIONS: Fragility metrics are primarily mathematical reflections of standard trial parameters and are heavily driven by the P value. Direct comparisons between fragility metrics and lost to follow-up are statistically inappropriate because LTFI and R-LTFI routinely exceed FI and RFI, respectively, and should be avoided in future fragility-based studies.

CLINICAL RELEVANCE: Understanding that statistical fragility metrics are mathematical transformations of trial parameters, and not independent measures of robustness, can prevent trial misinterpretation. Additionally, recognizing that the number of patients required to be lost to follow-up to reverse trial significance often exceeds fragility metrics should discourage inappropriate comparisons in future orthopaedic research.

Featured in the 2026-08-31 issue.

← Teaching Surgery Through Immediate Preoperative Small-Group D…Classifications in Brief: Kellgren-Lawrence Classification of… →

The Reduction is a free email digest of newly published orthopaedic literature — a handful of new papers in the subspecialties you choose, each summarized like this one. Subscribe free or browse the archive.