Introduction
Hysterectomy remains one of the most commonly performed gynecologic surgeries worldwide, with evolving techniques enhancing patient recovery and outcomes. Minimally invasive approaches—namely laparoscopic and robotic-assisted hysterectomy—have increasingly supplanted open surgery over recent decades, offering benefits such as reduced blood loss, shorter hospital stays, and faster convalescence. However, surgical proficiency is crucial, as technical performance directly influences perioperative morbidity and overall outcomes. Despite this, standardized, objective methods to assess surgeon skill in minimally invasive gynecologic surgery are limited, hampering quality assurance and training.
Development of the Assessment Tool
To address this gap, Tesfai et al. undertook an international, multi-center mixed-methods study to develop a comprehensive, objective assessment tool—the STELLAR score—for robotic-assisted and laparoscopic hysterectomy. Initial steps included a systematic literature review and the assembly of a steering group comprising seven expert gynecologists. Using Delphi consensus with 17 international experts from six countries, key procedural phases and potential technical errors or near misses in hysterectomy were identified.
The consensus-defined tool encapsulates seven distinct procedural phases alongside four quality metrics capturing surgical technique nuances, aiming to objectively quantify operative performance. This tool was designed to move beyond subjective global rating scales and conventional error assessment methods by incorporating structured phases and detailed quality indicators relevant to minimally invasive hysterectomy.
Validation and Reliability
Validation employed 40 unedited videos of minimally invasive hysterectomies assessed by multiple blinded raters in a national prospective multi-center observational study. Rigorous statistical analysis revealed excellent inter-rater and intra-rater reliability: intraclass correlation coefficients (ICC) of 0.969 and 0.810, respectively, indicating high agreement between different raters and consistent scoring over time. Internal consistency measured by Cronbach’s alpha across phases ranged from 0.743 to 0.834, supporting the homogeneity of the tool’s components.
Concurrent validity was confirmed through significant correlations with the established Observational Clinical Human Reliability Analysis (OCHRA) error assessment tool. This indicated the STELLAR score’s ability to accurately reflect known error rates in surgical performance.
Correlation with Clinical Outcomes
Crucially, the study demonstrated predictive validity by linking higher STELLAR scores—reflecting superior surgical technique—with better perioperative outcomes. Adjusting for confounding variables, elevated tool scores correlated with significantly shorter operative times and reduced blood loss. Moreover, higher scores were associated with a substantially lower incidence of postoperative complications graded by the Clavien-Dindo classification (Spearman’s rs = -0.438, p=0.004).
These findings underscore not only the tool’s measurement aptitude for surgical skill but also its clinical relevance. The ability to predict morbidity risk based on intraoperative performance highlights its potential use in surgical quality assurance, credentialing, and targeted training.
Clinical and Educational Implications
The STELLAR assessment tool offers a robust framework for objectively evaluating laparoscopic and robotic-assisted hysterectomy skills. By enabling granular phase-based feedback and quality measures, educators and supervisors can tailor interventions to specific technical deficits, reinforcing strengths and addressing weaknesses in trainees.
Furthermore, this validated instrument may serve as a benchmark in surgical trials assessing novel techniques or devices, ensuring consistent and reproducible performance standards. Institutions can incorporate it into credentialing pathways, promoting patient safety and continuous professional development.
Limitations and Future Directions
While the study benefits from international expert consensus and rigorous validation, limitations include reliance on video-based assessments, which may miss intraoperative contextual factors such as tactile feedback or team dynamics. Additionally, broader validation in diverse practice settings, including community hospitals and among less experienced surgeons, will further establish generalizability.
Future research should explore integration of this tool with emerging technologies such as machine learning-based video analysis for automated skill assessment. Longitudinal studies correlating assessment scores with long-term patient outcomes may deepen understanding of surgical proficiency impact.
Conclusion
This study solidifies the feasibility and validity of objectively measuring surgical performance in minimally invasive hysterectomy through a carefully developed and clinically relevant assessment tool. By linking technical proficiency with meaningful patient outcomes, the STELLAR score paves the way for standardized quality assurance, enhanced surgical education, and ultimately improved gynecologic care.
Reference
Tesfai FM, Iacobelli V, Chandrasekaran D, Lanceley A, Isabel MA, Preshaw J, Bala D, Pavone M, Bizzarri N, Rosati A, Mascagni P, Padoy N, Arboit L, Patel H, Seracchioli R, Saridogan E, Mabrouk M, Querleu D, Di Donato V, Lecointre L, Ahmed J, Vashisht A, Raimondo D, Capasso I, Vargiu V, Casanova J, Taliente F, Stoyanov D, Fagotti A, Francis N. Development and clinical validation of an objective assessment tool for robotic-assisted and laparoscopic hysterectomy. Am J Obstet Gynecol. 2026 Aug 12:S0002-9378(26)00413-8. doi: 10.1016/j.ajog.2026.08.010. Epub ahead of print. PMID: 42586187.

