A Pathologist-Annotated Dataset for Validating Artificial Intelligence: A Project Description and Pilot Study.

A Pathologist-Annotated Dataset for Validating Artificial Intelligence: A Project Description and Pilot Study.
复制标题

DOI:
10.4103/jpi.jpi_83_20
复制
发表时间:
2021
影响因子:
--
通讯作者:
Gallas BD
Gallas BD
中科院分区:
其他
文献类型:
--
作者:
Dudgeon SN;Wen S;Hanna MG;Gupta R;Amgad M;Sheth M;Marble H;Huang R;Herrmann MD;Szu CH;Tong D;Werness B;Szu E;Larsimont D;Madabhushi A;Hytopoulos E;Chen W;Singh R;Hart SN;Sharma A;Saltz J;Salgado R;Gallas BD

文献摘要

被引文献

相似文献

由于缺乏标准参考数据(地面实况),验证人工智能算法在医学图像中的临床应用是一项具有挑战性的工作。该主题通常只占研究论文讨论的一小部分,因为大部分工作都集中在开发新颖的算法上。在这项工作中,我们提出了一项合作,为处理整个幻灯片图像的算法创建病理学家注释的验证数据集。我们重点关注在估计乳腺癌中间质肿瘤浸润淋巴细胞 (sTIL) 密度的背景下的数据收集和算法性能评估。我们对在单一临床地点制备的 64 张苏木精和伊红染色的浸润性导管癌核心活检载玻片进行了数字化。合作的病理学家为每张玻片选择 10 个感兴趣区域 (ROI) 进行评估。我们创建了培训材料和工作流程,以两种模式众包病理学家图像注释:光学显微镜和两个数字平台。显微镜平台允许在两种模式下评估相同的 ROI。工作流程收集 ROI 类型、关于 ROI 是否适合估计 sTIL 密度的决定,以及该 ROI 的 sTIL 密度值(如果适用)。在一次数据收集活动和接下来的两周内,19 名病理学家总共进行了 1645 次 ROI 评估。初步研究产生了大量具有名义 sTIL 浸润的病例。此外,我们发现 sTIL 密度在一个病例内是相关的,并且存在显着的病理学家变异性。因此,我们概述了提高投资回报率和案例抽样方法的计划。我们还概述了统计方法,以在验证算法时考虑病例内的 ROI 相关性和病理学家的变异性。我们构建了高效数据收集的工作流程,并在试点研究中对其进行了测试。当我们准备关键研究时,我们将研究使用数据集作为算法的外部验证工具的方法。我们还将考虑如何使数据集适合监管目的:研究规模、患者群体以及病理学家培训和资格。为此,我们将通过医疗器械开发工具计划征求美国食品和药物管理局以及更广泛的数字病理学和人工智能社区的反馈。最终,我们打算分享数据集、统计方法和经验教训。
Validating artificial intelligence algorithms for clinical use in medical images is a challenging endeavor due to a lack of standard reference data (ground truth). This topic typically occupies a small portion of the discussion in research papers since most of the efforts are focused on developing novel algorithms. In this work, we present a collaboration to create a validation dataset of pathologist annotations for algorithms that process whole slide images. We focus on data collection and evaluation of algorithm performance in the context of estimating the density of stromal tumor-infiltrating lymphocytes (sTILs) in breast cancer. We digitized 64 glass slides of hematoxylin- and eosin-stained invasive ductal carcinoma core biopsies prepared at a single clinical site. A collaborating pathologist selected 10 regions of interest (ROIs) per slide for evaluation. We created training materials and workflows to crowdsource pathologist image annotations on two modes: an optical microscope and two digital platforms. The microscope platform allows the same ROIs to be evaluated in both modes. The workflows collect the ROI type, a decision on whether the ROI is appropriate for estimating the density of sTILs, and if appropriate, the sTIL density value for that ROI. In total, 19 pathologists made 1645 ROI evaluations during a data collection event and the following 2 weeks. The pilot study yielded an abundant number of cases with nominal sTIL infiltration. Furthermore, we found that the sTIL densities are correlated within a case, and there is notable pathologist variability. Consequently, we outline plans to improve our ROI and case sampling methods. We also outline statistical methods to account for ROI correlations within a case and pathologist variability when validating an algorithm. We have built workflows for efficient data collection and tested them in a pilot study. As we prepare for pivotal studies, we will investigate methods to use the dataset as an external validation tool for algorithms. We will also consider what it will take for the dataset to be fit for a regulatory purpose: study size, patient population, and pathologist training and qualifications. To this end, we will elicit feedback from the Food and Drug Administration via the Medical Device Development Tool program and from the broader digital pathology and AI community. Ultimately, we intend to share the dataset, statistical methods, and lessons learned.