External validation of Global Evaluative Assessment of Robotic Skills (GEARS)

External validation of Global Evaluative Assessment of Robotic Skills (GEARS)
复制标题

DOI:
10.1007/s00464-015-4070-8
复制
发表时间:
2015-11-01
影响因子:
3.1
通讯作者:
Goh, Alvin C.
Goh, Alvin C.
中科院分区:
医学2区
文献类型:
--
作者:
Aghazadeh, Monty A.;Jayaratna, Isuru S.;Goh, Alvin C.

文献摘要

被引文献

相似文献

我们在一个独立的队列中使用动物体内训练模型,证明了机器人技能全球评估评估(GEARS)的有效性、可靠性和实用性。GEARS是一种旨在衡量机器人技术技能的临床评估工具。采用横断面观察性研究设计,47名自愿参与者被分类为专家(bb30例机器人病例作为初级外科医生完成)或实习生。学员组进一步分为中级(千分之一日元5例,但千分之一货币符号30例)或新手(< 5例)。所有参与者都在猪模型中完成了一项标准化的体内机器人任务。任务表现由两名专家机器人外科医生评估,并由参与者使用GEARS评估工具进行自我评估。采用Kruskal-Wallis检验比较GEARS绩效得分,确定构念效度;Spearman等级相关测量观察者间信度;采用Cronbach’s alpha评价内部一致性。对9名专家和38名学员(中级14名,新手24名)完成绩效评估。与中级和新手相比,专家在整体和所有单独领域表现出更好的表现(p < 0.0001)。中级和新手在整体绩效上有显著性差异(p = 0.0505),而在效率和自主性的个体领域上有显著性差异(p = 0.0280和0.0425)。专家评分之间的观察者间信度被证实为强相关性(r = 0.857, 95% CI[0.691, 0.941])。专家和参与者评分的一致性较差(r = 0.435, 95% CI[0.121, 0.689]和r = 0.422, 95% CI[0.081, 0.0672])。专家和参与者的内部一致性很好(alpha = 0.96, 0.98, 0.93)。在一个独立的队列中,GEARS能够区分不同的机器人技能水平,证明了良好的构念效度。作为一种标准化的评估工具,GEARS在机器人体内手术任务中保持了一致性和可靠性,可以应用于广泛的机器人手术过程中的技能评估。
We demonstrate the construct validity, reliability, and utility of Global Evaluative Assessment of Robotic Skills (GEARS), a clinical assessment tool designed to measure robotic technical skills, in an independent cohort using an in vivo animal training model.Using a cross-sectional observational study design, 47 voluntary participants were categorized as experts (> 30 robotic cases completed as primary surgeon) or trainees. The trainee group was further divided into intermediates (a parts per thousand yen5 but a parts per thousand currency sign30 cases) or novices (< 5 cases). All participants completed a standardized in vivo robotic task in a porcine model. Task performance was evaluated by two expert robotic surgeons and self-assessed by the participants using the GEARS assessment tool. Kruskal-Wallis test was used to compare the GEARS performance scores to determine construct validity; Spearman's rank correlation measured interobserver reliability; and Cronbach's alpha was used to assess internal consistency.Performance evaluations were completed on nine experts and 38 trainees (14 intermediate, 24 novice). Experts demonstrated superior performance compared to intermediates and novices overall and in all individual domains (p < 0.0001). In comparing intermediates and novices, the overall performance difference trended toward significance (p = 0.0505), while the individual domains of efficiency and autonomy were significantly different between groups (p = 0.0280 and 0.0425, respectively). Interobserver reliability between expert ratings was confirmed with a strong correlation observed (r = 0.857, 95 % CI [0.691, 0.941]). Experts and participant scoring showed less agreement (r = 0.435, 95 % CI [0.121, 0.689] and r = 0.422, 95 % CI [0.081, 0.0672]). Internal consistency was excellent for experts and participants (alpha = 0.96, 0.98, 0.93).In an independent cohort, GEARS was able to differentiate between different robotic skill levels, demonstrating excellent construct validity. As a standardized assessment tool, GEARS maintained consistency and reliability for an in vivo robotic surgical task and may be applied for skills evaluation in a broad range of robotic procedures.