A Dataset and Benchmarks for Segmentation and Recognition of Gestures in Robotic Surgery.

A Dataset and Benchmarks for Segmentation and Recognition of Gestures in Robotic Surgery.
复制标题

DOI:
10.1109/tbme.2016.2647680
复制
发表时间:
2017-09
期刊:
IEEE transactions on bio-medical engineering
影响因子:
--
通讯作者:
Hager GD
Hager GD
中科院分区:
其他
文献类型:
--
作者:
Ahmidi N;Tao L;Sefati S;Gao Y;Lea C;Haro BB;Zappella L;Khudanpur S;Vidal R;Hager GD

文献摘要

被引文献

相似文献

最先进的手术数据分析技术报告了自动化技能评估和动作识别的有希望的结果。然而,许多这些技术的贡献仅限于研究特定的数据和验证指标,使得整个领域的进展评估极具挑战性。在本文中,我们解决了手术数据分析的两个主要问题:(1)缺乏统一的共享数据集和基准,(2)缺乏一致的验证过程。我们通过介绍JHU-ISI手势和技能评估工作集(JIGSAWS)来解决前者,这是我们创建的一个公共数据集,用于支持比较研究基准。JIGSAWS包含来自不同技能的操作者的机器人手术任务的多个性能的同步视频和运动学数据。我们解决后者提出了一个有据可查的评估方法和报告结果的六种技术的自动分割和分类的时间序列数据JIGSAWS。这些技术包括四种用于联合分割和分类的时间方法:隐马尔可夫模型,稀疏HMM,马尔可夫半马尔可夫条件随机场和跳跃链CRF;以及两种旨在对固定段进行分类的基于特征的方法:时空特征袋和线性动态系统。大多数方法识别手势活动,在留一超级试验和留一用户交叉验证设置下的总体准确率约为80%。目前的方法在这个共享数据集上显示出有希望的结果,但仍有取得重大进展的空间,特别是在不同外科医生之间一致预测手势活动方面。本文报告的结果提供了第一个系统和统一的评估基准数据库上的手术活动识别技术。
State-of-the-art techniques for surgical data analysis report promising results for automated skill assessment and action recognition. The contributions of many of these techniques, however, are limited to study-specific data and validation metrics, making assessment of progress across the field extremely challenging. In this paper, we address two major problems for surgical data analysis: (1) lack of uniform shared datasets and benchmarks and (2) lack of consistent validation processes. We address the former by presenting the JHU-ISI Gesture and Skill Assessment Working Set (JIGSAWS), a public dataset we have created to support comparative research benchmarking. JIGSAWS contains synchronized video and kinematic data from multiple performances of robotic surgical tasks by operators of varying skill. We address the latter by presenting a well-documented evaluation methodology and reporting results for six techniques for automated segmentation and classification of time-series data on JIGSAWS. These techniques comprise four temporal approaches for joint segmentation and classification: Hidden Markov Model, Sparse HMM, Markov semi-Markov Conditional Random Field, and Skip-Chain CRF; and two feature-based ones that aim to classify fixed segments: Bag of spatiotemporal Features and Linear Dynamical Systems. Most methods recognize gesture activities with approximately 80% overall accuracy under both leave-one-super-trial-out and leave-one-user-out cross-validation settings. Current methods show promising results on this shared dataset, but room for significant progress remains, particularly for consistent prediction of gesture activities across different surgeons. The results reported in this paper provide the first systematic and uniform evaluation of surgical activity recognition techniques on the benchmark database.