Measuring Spine Surgeon Performance: A Scoping Review of Assessment Metrics and Evaluation Methods Used to Assess Surgeon Competency.

Spine (Phila Pa 1976) · Oct 15 2025 · Review

McMahan SR, Courtois EC, Satin AM, Guyer RD, Wilson BA, Robinson KT, et al.

Texas Back Institute, Plano, TX

Spine

SUMMARY — THE REDUCTIONThis scoping review of 44 studies finds spine surgery training assessments heavily emphasize technical skills with validated tools, while nontechnical skills remain inconsistently and inadequately measured.
Abstract, as published

OBJECTIVE: To systematically review and synthesize the performance metrics used to assess surgical competency during spine surgery training.

SUMMARY OF BACKGROUND DATA: The complexity of spine surgery requires prolonged training and careful competency evaluation. Simulation-based training offers scalable, repeatable, and ethically feasible alternatives to cadaver-based education. However, the assessment of surgeon performance across these platforms varies widely in scope and standardization.

METHODS: A scoping review was conducted following the PRISMA-ScR guidelines and registered in the Open Science Framework. Three databases were searched through February 2025 for prospective studies assessing surgeon performance in spine surgery training. Included studies evaluated technical and/or nontechnical skills using defined metrics across various simulation platforms. Data extraction focused on surgical procedure, simulator type, assessment metrics, and scoring methods.

RESULTS: From 974 screened records, 44 studies were included. Technical skills (TS) were assessed in all studies, primarily focusing on accuracy, efficiency, handling, safety, and efficacy. Nontechnical skills (NTS)-including cognition, communication, and self-assessment-were reported in 12 studies. Assessment metrics were influenced by surgical procedure and simulation modality. Physical models were most frequently used (n=25), followed by virtual (n=8), hybrid (n=9), and cadaveric or patient models. Scoring systems ranged from validated tools (eg, OSATS, GRS) to piloted instruments. TS were often measured via reviewer scoring or automated simulator output, while NTS assessments lacked consistency and standardization.

CONCLUSION: Performance assessments in spine surgery simulation training vary significantly across platforms and procedures. TS are widely measured using objective or structured scoring systems, whereas NTS remain underassessed. This review underscores the need for validated, comprehensive, and procedure-specific performance metrics-integrating both TS and NTS-to enhance training, standardize evaluation, and ensure clinical readiness.

STUDY DESIGN: This is a scoping review.

Featured in the 2026-08-10 issue.

← The SCARE Score: A Simplified Risk Score for Predicting Intra…Geometric foundations of measurement uncertainty and clinical… →

The Reduction is a free email digest of newly published orthopaedic literature — a handful of new papers in the subspecialties you choose, each summarized like this one. Subscribe free or browse the archive.