Detection of frame informativeness in endoscopic videos using image quality and recurrent neural networks

Detection of frame informativeness in endoscopic videos using image quality and recurrent neural networks
复制标题

使用图像质量和循环神经网络检测内窥镜视频中的帧信息

DOI:
10.1117/12.2545734
复制
发表时间:
2020
期刊:
--
影响因子:
--
通讯作者:
P. D. With
P. D. With
中科院分区:
--
文献类型:
--
作者:
T. Boers;J. V. D. Putten;J. D. Groof;M. Struyvenberg;K. Fockens;W. Curvers;E. Schoon;F. V. D. Sommen;J. Bergman;P. D. With

文献摘要

被引文献

相似文献

据估计,胃肠病学家在巴雷特食管患者中误诊高达25%的食管腺癌。这提示需要更敏感和客观的工具来帮助临床医生进行病变检测。人工智能(AI)可以使考试更加客观,因此有助于减轻对观察者的依赖。由于这些模型使用高质量的内窥镜视频帧进行训练以获得高效率,因此还需要高质量的图像进行推理。因此,我们的目标是开发一个框架,能够区分良好的图像质量的先验信息分类,从而导致高推理鲁棒性。我们表明,我们可以使用递归神经网络在时间域上保持信息量,与对单个图像进行分类相比,在非信息量检测方面具有更高的性能。此外,还发现,通过使用梯度加权类激活图(Grad-CAM),我们可以更好地定位帧内的信息。我们已经开发了一个定制的Resnet18特征提取器,其中包含3个分类器,包括全连接(FC),长短期记忆(LSTM)和门控递归单元(GRU)分类器。实验结果基于来自食道的20个拉回视频的4,349帧。我们的结果表明,该算法实现了与当前最先进的性能相当的性能。FC和LSTM分类器达到了91%和91%的F1分数。我们发现,基于LSTM分类器的Grad-CAM最能代表非信息性的来源,因为85%的图像被发现突出显示了正确的区域。我们用于内窥镜信息性分类的新实现的益处在于,它是端到端训练的,在决策中结合了时空域以实现鲁棒性,并且使用Grad-CAM使模型的模型决策具有洞察力。
Gastroenterologists are estimated to misdiagnose up to 25% of esophageal adenocarcinomas in Barrett's Esophagus patients. This prompts the need for more sensitive and objective tools to aid clinicians with lesion detection. Artificial Intelligence (AI) can make examinations more objective and will therefore help to mitigate the observer dependency. Since these models are trained with good-quality endoscopic video frames to attain high efficacy, high-quality images are also needed for inference. Therefore, we aim to develop a framework that is able to distinguish good image quality by a-priori informativeness classification which leads to high inference robustness. We show that we can maintain informativeness over the temporal domain using recurrent neural networks, yielding a higher performance on non-informativeness detection compared to classifying individual images. Furthermore, it is also found that by using Gradient weighted Class Activation Map (Grad-CAM), we can better localize informativeness within a frame. We have developed a customized Resnet18 feature extractor with 3 classifiers, consisting of a Fully-Connected (FC), Long-Short-Term-Memory (LSTM) and a Gated-Recurrent-Unit (GRU) classifier. Experimental results are based on 4,349 frames from 20 pullback videos of the esophagus. Our results demonstrate that the algorithm achieves comparative performance with the current state-of-the-art. The FC and LSTM classifier reach an F1 score of 91% and 91%. We found that the LSTM classifier based Grad-CAMs represent the origin of non-informativeness the best as 85% of the images were found to be highlighting the correct area. The benefit of our novel implementation for endoscopic informativeness classification is that it is trained end- to-end, incorporates the spatiotemporal domain in the decision making for robustness, and makes the model decisions of the model insightful with the use of Grad-CAMs.