Automatic Detection of the Pharyngeal Phase in Raw Videos for the Videofluoroscopic Swallowing Study Using Efficient Data Collection and 3D Convolutional Networks

Automatic Detection of the Pharyngeal Phase in Raw Videos for the Videofluoroscopic Swallowing Study Using Efficient Data Collection and 3D Convolutional Networks
复制标题

DOI:
10.3390/s19183873
复制
发表时间:
2019-09-02
期刊:
影响因子:
3.9
通讯作者:
Jung, Tae-Du
Jung, Tae-Du
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Lee, Jong Taek;Park, Eunhee;Jung, Tae-Du

文献摘要

被引文献

相似文献

视频荧光吞咽检查(VFSS)是吞咽困难的标准诊断工具。为了检测吞咽期间是否存在抽吸,通常使用手动搜索来在相应的VFSS图像上标记咽期的时间间隔。在这项研究中,我们提出了一种新的方法,使用3D卷积网络来检测原始VFSS视频中的咽部相位,而无需手动注释。为了有效地收集训练数据,我们提出了一个级联框架,该框架不再需要吞咽过程的时间间隔,也不需要手动标记用于检测的解剖位置。对于视频分类,我们应用了膨胀的3D卷积网络(I3D),这是最先进的动作分类网络之一,作为基线架构。我们还提出了一个修改后的3D卷积网络架构,该架构源自基线I3D架构。这两种架构的分类和检测性能进行了评估比较。实验结果表明,在两种模型都使用随机权重进行训练的情况下,该模型的性能优于基线I3D模型。我们的结论是,该方法大大减少了VFSS图像的检查时间与低失误率。
Videofluoroscopic swallowing study (VFSS) is a standard diagnostic tool for dysphagia. To detect the presence of aspiration during a swallow, a manual search is commonly used to mark the time intervals of the pharyngeal phase on the corresponding VFSS image. In this study, we present a novel approach that uses 3D convolutional networks to detect the pharyngeal phase in raw VFSS videos without manual annotations. For efficient collection of training data, we propose a cascade framework which no longer requires time intervals of the swallowing process nor the manual marking of anatomical positions for detection. For video classification, we applied the inflated 3D convolutional network (I3D), one of the state-of-the-art network for action classification, as a baseline architecture. We also present a modified 3D convolutional network architecture that is derived from the baseline I3D architecture. The classification and detection performance of these two architectures were evaluated for comparison. The experimental results show that the proposed model outperformed the baseline I3D model in the condition where both models are trained with random weights. We conclude that the proposed method greatly reduces the examination time of the VFSS images with a low miss rate.