Attention-Guided 3D-CNN Framework for Glaucoma Detection and Structural-Functional Association Using Volumetric Images.

Attention-Guided 3D-CNN Framework for Glaucoma Detection and Structural-Functional Association Using Volumetric Images.
复制标题

DOI:
10.1109/jbhi.2020.3001019
复制
发表时间:
2020-12
影响因子:
7.7
通讯作者:
Garnavi R
Garnavi R
中科院分区:
工程技术1区
文献类型:
--
作者:
George Y;Antony BJ;Ishikawa H;Wollstein G;Schuman JS;Garnavi R

文献摘要

被引文献

相似文献

对3D光学相干断层扫描(OCT)体积的直接分析使深度学习模型(DL)能够学习空间结构信息并发现与青光眼相关的新生物标志物。下采样3D输入体积是最先进的解决方案,以适应有限数量的训练体积以及可用的计算资源。然而,这限制了网络从OCT体积中的小视网膜结构学习的能力。在本文中,我们的目标是通过在训练过程中为DL模型提供指导来提高性能,以便从3D OCT体积中更精细的眼部结构中学习。因此,我们提出了一个端到端的注意力引导的3D DL模型,用于青光眼检测和从视网膜结构估计视觉功能。该模型由三个路径组成,具有相同的网络结构,但不同的输入。一个输入是原始3D-OCT立方体,另外两个是在3D梯度类激活热图引导的训练期间计算的。每个路径输出类别标签,整个模型同时进行训练,以最大限度地减少三个路径的损失之和。通过融合三个路径的预测获得最终输出。此外,为了探索所提出的模型的鲁棒性和可推广性,我们将该模型应用于青光眼检测的分类任务以及估计视野指数(VFI)(0至100之间的值)的回归任务。使用总共3782和10,370次OCT扫描的5重交叉验证分别训练和评估分类和回归模型。青光眼检测模型的曲线下面积(AUC)为93.8%,而没有注意力引导组件的基线模型为86.8%。该模型还优于六种不同的基于特征的机器学习方法,这些方法使用扫描仪计算的测量值进行训练。此外,我们还评估了与青光眼相关的不同视网膜层的贡献。VFI估计模型实现了皮尔逊相关性和绝对误差中位数分别为0.75和3.6%,为3100立方的测试集。
The direct analysis of 3D Optical Coherence Tomography (OCT) volumes enables deep learning models (DL) to learn spatial structural information and discover new bio-markers that are relevant to glaucoma. Down-sampling 3D input volumes is the state-of-art solution to accommodate for the limited number of training volumes as well as the available computing resources. However, this limits the network’s ability to learn from small retinal structures in OCT volumes. In this paper, our goal is to improve the performance by providing guidance to DL model during training in order to learn from finer ocular structures in 3D OCT volumes. Therefore, we propose an end-to-end attention guided 3D DL model for glaucoma detection and estimating visual function from retinal structures. The model consists of three pathways with the same network architecture but different inputs. One input is the original 3D-OCT cube and the other two are computed during training guided by the 3D gradient class activation heatmaps. Each pathway outputs the class-label and the whole model is trained concurrently to minimize the sum of losses from three pathways. The final output is obtained by fusing the predictions of the three pathways. Also, to explore the robustness and generalizability of the proposed model, we apply the model on a classification task for glaucoma detection as well as a regression task to estimate visual field index (VFI) (a value between 0 and 100). A 5-fold cross-validation with a total of 3782 and 10,370 OCT scans is used to train and evaluate the classification and regression models, respectively. The glaucoma detection model achieved an area under the curve (AUC) of 93.8% compared with 86.8% for a baseline model without the attention-guided component. The model also outperformed six different feature based machine learning approaches that use scanner computed measurements for training. Further, we also assessed the contribution of different retinal layers that are relevant to glaucoma. The VFI estimation model achieved a Pearson correlation and median absolute error of 0.75 and 3.6%, respectively, for a test set of size 3100 cubes.