Use of Crowd Innovation to Develop an Artificial Intelligence-Based Solution for Radiation Therapy Targeting

Use of Crowd Innovation to Develop an Artificial Intelligence-Based Solution for Radiation Therapy Targeting
复制标题

DOI:
10.1001/jamaoncol.2019.0159
复制
发表时间:
2019-05-01
期刊:
影响因子:
28.4
通讯作者:
Guinan, Eva C.
Guinan, Eva C.
中科院分区:
医学1区
文献类型:
--
作者:
Mak, Raymond H.;Endres, Michael G.;Guinan, Eva C.

文献摘要

被引文献

相似文献

放射治疗(RT)是一种重要的癌症治疗方法,但现有的放射肿瘤学家工作队伍无法满足日益增长的全球需求。在放疗计划中,医生的一项关键任务是肿瘤的靶向分割,这需要大量的训练,并受到观察者之间的显著差异的影响。目的确定群体创新是否可以用于快速生产人工智能(Al)解决方案,这些解决方案可以复制放射肿瘤学专家在肺肿瘤分割RT靶向方面的准确性。设置。我们进行了一个为期10周、基于奖金的在线挑战,分为3个阶段(奖金总额为55美元)。比赛使用了一个精心策划的数据集,包括计算机断层扫描(CT)扫描和临床护理专家生成的肺肿瘤分割(CT扫描来自461名患者;平均每次扫描157张图像;总共77942张图像;8144张存在肿瘤的图像)。参赛者被提供了229个CT扫描的训练集,并附带专家轮廓来开发他们的算法,并在整个比赛过程中获得他们的表现反馈,包括来自专家临床医生的反馈。参赛者生成的人工智能算法在一个独立的数据集上自动评分,该数据集不向参赛者提供,并使用定量指标评估每个算法的自动分割与专家分割的重叠程度,对性能进行排名。性能进一步针对人类专家观察者之间和观察者内部的变化进行基准测试。结果来自62个国家的564名参赛者报名参加了本次挑战赛,其中34名(6%)提交了算法。由前5名人工智能算法产生的自动分割,当与集成模型相结合时,其精度(Dice系数= 0.79)在6名人类专家之间测量的平均观察者间变化的基准范围内。对于第一阶段,排名前7的算法在holdout数据集上的平均自定义分割分数(5分)在0.15到0.38之间,使用相对误差度量的性能不是最优的。第二阶段的平均得分增加到0.53到037,其他性能指标也有类似的改善。在阶段3中,顶级算法的性能又提高了9%。使用集成模型将阶段2和阶段3的前5种算法结合起来,性能提高了9%到12%,最终得分达到0.68。人群创新和人工智能相结合的方法迅速产生了自动化算法,复制了训练有素的医生在放射治疗关键任务中的技能。这些人工智能算法可以通过将专家临床医生的技能转移到资源不足的医疗机构,从而改善全球的癌症护理。
IMPORTANCE Radiation therapy (RT) is a critical cancer treatment, but the existing radiation oncologist work force does not meet growing global demand. One key physician task in RT planning involves tumor segmentation for targeting, which requires substantial training and is subject to significant interobserver variation.OBJECTIVE To determine whether crowd innovation could be used to rapidly produce artificial intelligence (Al) solutions that replicate the accuracy of an expert radiation oncologist in segmenting lung tumors for RT targeting.DESIGN. SETTING. AND PARTICIPANTS We conducted a 10-week, prize-based, online, 3-phase challenge (prizes totaled $55 0 0 0). A well-curated data set, including computed tomographic (CT) scans and lung tumor segmentations generated by an expert for clinical care, was used for the contest (CT scans from 461 patients; median 157 images per scan; 77 942 images in total; 8144 images with tumor present). Contestants were provided a training set of 229 CT scans with accompanying expert contours to develop their algorithms and given feedback on their performance throughout the contest, including from the expert clinician.MAIN OUTCOMES AND MEASURES The Al algorithms generated by contestants were automatically scored on an independent data set that was withheld from contestants, and performance ranked using quantitative metrics that evaluated overlap of each algorithm's automated segmentations with the expert's segmentations. Performance was further benchmarked against human expert interobserver and intraobserver variation.RESULTS A total of 564 contestants from 62 countries registered for this challenge, and 34 (6%) submitted algorithms. The automated segmentations produced by the top 5 Al algorithms, when combined using an ensemble model, had an accuracy (Dice coefficient = 0.79) that was within the benchmark of mean interobserver variation measured between 6 human experts. For phase 1, the top 7 algorithms had average custom segmentation scores (5 scores) on the holdout data set ranging from 0.15 to 0.38, and suboptimal performance using relative measures of error. The average scores for phase 2 increased to 0.53 to 037, with a similar improvement in other performance metrics. In phase 3, performance of the top algorithm increased by an additional 9%. Combining the top 5 algorithms from phase 2 and phase 3 using an ensemble model, yielded an additional 9% to 12% improvement in performance with a final score reaching 0.68.CONCLUSIONS ANC RELEVANCE A combined crowd innovation and Al approach rapidly produced automated algorithms that replicated the skills of a highly trained physician for a critical task in radiation therapy. These Al algorithms could improve cancer care globally by transferring the skills of expert clinicians to under-resourced health care settings.