The intersection of video capsule endoscopy and artificial intelligence: addressing unique challenges using machine learning

The intersection of video capsule endoscopy and artificial intelligence: addressing unique challenges using machine learning
复制标题

DOI:
10.48550/arxiv.2308.13035
复制
发表时间:
2023-08
期刊:
ArXiv
影响因子:
--
通讯作者:
S. Guleria;B. Schwartz;Yash Sharma;Philip Fernandez;James A. Jablonski;Sodiq Adewole;S. Srivastava;Fisher Rhoads;Michael D. Porter;Michelle Yeghyayan;Dylan M. Hyatt;Andrew Copland;L. Ehsan;Donald E. Brown;S. Syed
S. Guleria;B. Schwartz;Yash Sharma;Philip Fernandez;James A. Jablonski;Sodiq Adewole;S. Srivastava;Fisher Rhoads;Michael D. Porter;Michelle Yeghyayan;Dylan M. Hyatt;Andrew Copland;L. Ehsan;Donald E. Brown;S. Syed
中科院分区:
其他
文献类型:
--
作者:
S. Guleria;B. Schwartz;Yash Sharma;Philip Fernandez;James A. Jablonski;Sodiq Adewole;S. Srivastava;Fisher Rhoads;Michael D. Porter;Michelle Yeghyayan;Dylan M. Hyatt;Andrew Copland;L. Ehsan;Donald E. Brown;S. Syed

文献摘要

相似文献

引言:技术负担和时间密集型审查过程限制了视频胶囊式内窥镜(VCE)的实际应用。人工智能(AI)有望解决这些限制,但AI和VCE的交叉揭示了必须首先克服的挑战。我们确定了要应对的五项挑战。挑战1:VCE数据是随机的,包含显著伪影。挑战#2:VCE解释成本高昂。挑战3:VCE数据本身就不平衡。挑战4:现有VCE AIMLT计算繁琐。挑战5:临床医生不愿接受无法解释其过程的AIMLT。研究方法:解剖标志检测模型用于测试卷积神经网络(CNN)在VCE数据分类任务中的应用。我们还创建了一个工具,帮助专家注释VCE数据。然后,我们使用不同的方法创建了更精细的模型,包括多帧方法,基于图形表示的CNN和基于元学习的少量方法。结果如下:当用于全长VCE镜头时,CNN准确地识别了解剖标志(99.1%),梯度加权类激活映射显示了CNN用于做出决定的每帧部分。具有弱监督学习的图CNN(准确率89.9%,灵敏度91.1%),少镜头模型(准确率90.8%,精度91.4%,灵敏度90.9%)和多帧模型(准确率97.5%,精度91.5%,灵敏度94.8%)表现良好。讨论:这五个挑战中的每一个都在一定程度上由我们的一个基于人工智能的模型来解决。我们的目标是使用旨在提高临床医生信心的轻量级模型来实现高性能。
Introduction: Technical burdens and time-intensive review processes limit the practical utility of video capsule endoscopy (VCE). Artificial intelligence (AI) is poised to address these limitations, but the intersection of AI and VCE reveals challenges that must first be overcome. We identified five challenges to address. Challenge #1: VCE data are stochastic and contains significant artifact. Challenge #2: VCE interpretation is cost-intensive. Challenge #3: VCE data are inherently imbalanced. Challenge #4: Existing VCE AIMLT are computationally cumbersome. Challenge #5: Clinicians are hesitant to accept AIMLT that cannot explain their process. Methods: An anatomic landmark detection model was used to test the application of convolutional neural networks (CNNs) to the task of classifying VCE data. We also created a tool that assists in expert annotation of VCE data. We then created more elaborate models using different approaches including a multi-frame approach, a CNN based on graph representation, and a few-shot approach based on meta-learning. Results: When used on full-length VCE footage, CNNs accurately identified anatomic landmarks (99.1%), with gradient weighted-class activation mapping showing the parts of each frame that the CNN used to make its decision. The graph CNN with weakly supervised learning (accuracy 89.9%, sensitivity of 91.1%), the few-shot model (accuracy 90.8%, precision 91.4%, sensitivity 90.9%), and the multi-frame model (accuracy 97.5%, precision 91.5%, sensitivity 94.8%) performed well. Discussion: Each of these five challenges is addressed, in part, by one of our AI-based models. Our goal of producing high performance using lightweight models that aim to improve clinician confidence was achieved.