Key challenges for delivering clinical impact with artificial intelligence

Key challenges for delivering clinical impact with artificial intelligence
复制标题

DOI:
10.1186/s12916-019-1426-2
复制
发表时间:
2019-10-29
期刊:
影响因子:
9.3
通讯作者:
King, Dominic
King, Dominic
中科院分区:
医学1区
文献类型:
--
作者:
Kelly, Christopher J.;Karthikesalingam, Alan;King, Dominic

文献摘要

被引文献

相似文献

背景:医疗保健领域的人工智能(AI)研究正在迅速加速,医学的各个领域都展示了潜在的应用。然而,目前成功地将这种技术应用于临床的例子有限。本文探讨了人工智能在医疗保健领域的主要挑战和限制,并考虑了将这些潜在的变革性技术从研究转化为临床实践所需的步骤。主要内容:医疗保健领域人工智能系统翻译的关键挑战包括机器学习科学固有的挑战、实施中的后勤困难、对采用障碍的考虑以及必要的社会文化或路径变化。作为随机对照试验的一部分,可靠的同行评议临床评估应被视为证据生成的黄金标准,但在实践中进行这些评估可能并不总是合适或可行的。性能指标的目标应该是捕捉真正的临床适用性,并为目标用户所理解。需要在创新速度与潜在危害之间取得平衡的监管,以及深思熟虑的上市后监督,以确保患者不会接触到危险的干预措施,也不会被剥夺获得有益创新的机会。必须开发能够对人工智能系统进行直接比较的机制,包括使用独立的、本地的和具有代表性的测试集。人工智能算法的开发者必须警惕潜在的危险,包括数据集转移、混杂因素的意外拟合、无意的歧视性偏见、推广到新人群的挑战,以及新算法对健康结果的意想不到的负面后果。结论:将人工智能研究安全及时地转化为经过临床验证和适当监管的系统,使所有人都能受益,这是具有挑战性的。稳健的临床评估,使用对临床医生直观的衡量标准,理想情况下超越技术准确性的衡量标准,包括护理质量和患者结果,是必不可少的。还需要进一步的工作(1)识别算法偏差和不公平的主题,同时制定缓解措施来解决这些问题,(2)减少脆性并提高泛化能力,以及(3)开发改进机器学习预测的可解释性的方法。如果这些目标能够实现,对患者的好处可能是变革性的。
Background: Artificial intelligence (AI) research in healthcare is accelerating rapidly, with potential applications being demonstrated across various domains of medicine. However, there are currently limited examples of such techniques being successfully deployed into clinical practice. This article explores the main challenges and limitations of AI in healthcare, and considers the steps required to translate these potentially transformative technologies from research to clinical practice.Main body: Key challenges for the translation of AI systems in healthcare include those intrinsic to the science of machine learning, logistical difficulties in implementation, and consideration of the barriers to adoption as well as of the necessary sociocultural or pathway changes. Robust peer-reviewed clinical evaluation as part of randomised controlled trials should be viewed as the gold standard for evidence generation, but conducting these in practice may not always be appropriate or feasible. Performance metrics should aim to capture real clinical applicability and be understandable to intended users. Regulation that balances the pace of innovation with the potential for harm, alongside thoughtful post-market surveillance, is required to ensure that patients are not exposed to dangerous interventions nor deprived of access to beneficial innovations. Mechanisms to enable direct comparisons of AI systems must be developed, including the use of independent, local and representative test sets. Developers of AI algorithms must be vigilant to potential dangers, including dataset shift, accidental fitting of confounders, unintended discriminatory bias, the challenges of generalisation to new populations, and the unintended negative consequences of new algorithms on health outcomes.Conclusion: The safe and timely translation of AI research into clinically validated and appropriately regulated systems that can benefit everyone is challenging. Robust clinical evaluation, using metrics that are intuitive to clinicians and ideally go beyond measures of technical accuracy to include quality of care and patient outcomes, is essential. Further work is required (1) to identify themes of algorithmic bias and unfairness while developing mitigations to address these, (2) to reduce brittleness and improve generalisability, and (3) to develop methods for improved interpretability of machine learning predictions. If these goals can be achieved, the benefits for patients are likely to be transformational.