A multi-institutional study using artificial intelligence to provide reliable and fair feedback to surgeons.

A multi-institutional study using artificial intelligence to provide reliable and fair feedback to surgeons.
复制标题

DOI:
10.1038/s43856-023-00263-3
复制
发表时间:
2023-03-30
期刊:
COMMUNICATIONS MEDICINE
影响因子:
--
通讯作者:
Hung, Andrew J
Hung, Andrew J
中科院分区:
其他
文献类型:
--
作者:
Kiyasseh, Dani;Laca, Jasper;Haque, Taseen F;Miles, Brian J;Wagner, Christian;Donoho, Daniel A;Anandkumar, Animashree;Hung, Andrew J

文献摘要

被引文献

相似文献

外科医生如果能收到关于其表现的可靠反馈,就能迅速掌握手术所需的技能。这种基于表现的反馈可以由最近开发的人工智能(AI)系统提供,该系统基于手术视频评估外科医生的技能,同时突出显示与评估最相关的视频方面。然而,这些亮点或解释是否对所有外科医生都同样可靠仍然是一个悬而未决的问题。在这里,我们通过将来自两大洲三家医院的手术视频与人类专家生成的解释进行比较,系统地量化了基于人工智能的解释的可靠性。为了提高基于AI的解释的可靠性,我们提出了解释训练策略-TWIX -它使用人类解释作为监督,明确地教AI系统突出重要的视频帧。我们表明,虽然基于人工智能的解释通常与人类的解释一致,但它们对于不同的外科医生子队列(例如,新手对专家),我们称之为解释偏差的现象。我们还表明,TWIX增强了基于AI的解释的可靠性,减轻了解释偏差,并提高了医院AI系统的性能。这些发现延伸到一个培训环境中,医学生可以提供反馈今天。我们的研究为即将实施的人工智能增强手术培训和外科医生资格认证计划提供了信息,并有助于手术的安全和公平民主化。外科医生的目标是掌握外科手术所必需的技能。其中一项技能是通过一系列的针脚将物体连接在一起。掌握这些手术技能可以通过向外科医生提供关于其表现质量的反馈来提高。然而,这种反馈在外科实践中常常是不存在的。虽然理论上可以通过最近开发的人工智能(AI)系统提供基于性能的反馈,该系统使用计算模型来评估外科医生的技能,但这种反馈的可靠性仍然未知。在这里,我们将基于人工智能的反馈与人类专家提供的反馈进行比较,并证明它们经常相互重叠。我们还表明,明确地教导人工智能系统与人类反馈保持一致,进一步提高了基于人工智能的新手术视频反馈的可靠性。我们的研究结果概述了人工智能系统通过提供可靠且专注于特定技能的反馈来支持外科医生培训的潜力,并指导通过补充技能评估来提高此类评估可信度的程序。Kiyasseh等人比较了人工智能(AI)向外科医生提供的反馈质量与人类专家提供的反馈质量。教导人工智能系统明确遵循人类解释可以提高可靠性,并减少基于人工智能的反馈的偏差。
Surgeons who receive reliable feedback on their performance quickly master the skills necessary for surgery. Such performance-based feedback can be provided by a recently-developed artificial intelligence (AI) system that assesses a surgeon’s skills based on a surgical video while simultaneously highlighting aspects of the video most pertinent to the assessment. However, it remains an open question whether these highlights, or explanations, are equally reliable for all surgeons. Here, we systematically quantify the reliability of AI-based explanations on surgical videos from three hospitals across two continents by comparing them to explanations generated by humans experts. To improve the reliability of AI-based explanations, we propose the strategy of training with explanations –TWIX –which uses human explanations as supervision to explicitly teach an AI system to highlight important video frames. We show that while AI-based explanations often align with human explanations, they are not equally reliable for different sub-cohorts of surgeons (e.g., novices vs. experts), a phenomenon we refer to as an explanation bias. We also show that TWIX enhances the reliability of AI-based explanations, mitigates the explanation bias, and improves the performance of AI systems across hospitals. These findings extend to a training environment where medical students can be provided with feedback today. Our study informs the impending implementation of AI-augmented surgical training and surgeon credentialing programs, and contributes to the safe and fair democratization of surgery. Surgeons aim to master skills necessary for surgery. One such skill is suturing which involves connecting objects together through a series of stitches. Mastering these surgical skills can be improved by providing surgeons with feedback on the quality of their performance. However, such feedback is often absent from surgical practice. Although performance-based feedback can be provided, in theory, by recently-developed artificial intelligence (AI) systems that use a computational model to assess a surgeon’s skill, the reliability of this feedback remains unknown. Here, we compare AI-based feedback to that provided by human experts and demonstrate that they often overlap with one another. We also show that explicitly teaching an AI system to align with human feedback further improves the reliability of AI-based feedback on new videos of surgery. Our findings outline the potential of AI systems to support the training of surgeons by providing feedback that is reliable and focused on a particular skill, and guide programs that give surgeons qualifications by complementing skill assessments with explanations that increase the trustworthiness of such assessments. Kiyasseh et al. compare the quality of feedback provided to surgeons by artificial intelligence (AI) to that provided by human experts. Teaching an AI system to explicitly follow human explanations improves the reliability and reduces the bias of AI-based feedback.