Gaze-assisted automatic captioning of fetal ultrasound videos using three-way multi-modal deep neural networks

Gaze-assisted automatic captioning of fetal ultrasound videos using three-way multi-modal deep neural networks
复制标题

DOI:
10.1016/j.media.2022.102630
复制
发表时间:
2022-10-09
影响因子:
10.9
通讯作者:
Noble, J. Alison
Noble, J. Alison
中科院分区:
工程技术1区
文献类型:
--
作者:
Alsharid, Mohammad;Cai, Yifan;Noble, J. Alison

文献摘要

被引文献

相似文献

在这项工作中,我们提出了一种新的凝视辅助自然语言处理(NLP)为基础的视频字幕模型来描述常规的中期妊娠胎儿超声扫描视频中的口头超声检查的词汇。我们的多模态方法的主要新奇在于,学习视频字幕模型是使用超声波视频,跟踪凝视和语音记录的文本transanimation的组合构建的。描述时空扫描视频内容的文本字幕是从声谱仪语音记录中学习的。字幕的生成由超声波检查者注视跟踪信息辅助,该信息反映了他们在执行实时成像和解释冻结图像时的视觉注意力。为了评估在视频模型上添加或保留不同形式的凝视的效果,我们比较了使用三种多模态配置训练的时空深度网络,即:(1)仅具有文本和视频作为输入的无注视神经网络,(2)另外使用注意力图形式的真实的超声医师注视的神经网络,以及(3)替代地使用显著图形式的自动预测注视的神经网络。我们通过建立一般的基于文本的指标(BLEU,ROUGE-L,F1得分),特定于域的指标(ARS),并考虑生成的字幕的丰富性和效率的扫描视频的指标来评估算法的性能。结果表明,所提出的凝视辅助模型可以产生更丰富和更多样化的字幕的临床胎儿超声扫描视频比那些没有凝视的代价感知句子结构。结果还表明,所生成的字幕是类似的声谱仪语音讨论的视觉内容和扫描动作执行。
In this work, we present a novel gaze-assisted natural language processing (NLP)-based video captioning model to describe routine second-trimester fetal ultrasound scan videos in a vocabulary of spoken sonography. The primary novelty of our multi-modal approach is that the learned video captioning model is built using a combination of ultrasound video, tracked gaze and textual transcriptions from speech recordings. The textual captions that describe the spatio-temporal scan video content are learnt from sonographer speech recordings. The generation of captions is assisted by sonographer gaze-tracking information reflecting their visual attention while performing live-imaging and interpreting a frozen image. To evaluate the effect of adding, or withholding, different forms of gaze on the video model, we compare spatio-temporal deep networks trained using three multi-modal configurations, namely: (1) a gaze-less neural network with only text and video as input, (2) a neural network additionally using real sonographer gaze in the form of attention maps, and (3) a neural network using automatically-predicted gaze in the form of saliency maps instead. We assess algorithm performance through established general text-based metrics (BLEU, ROUGE-L, F1 score), a domain-specific metric (ARS), and metrics that consider the richness and efficiency of the generated captions with respect to the scan video. Results show that the proposed gaze-assisted models can generate richer and more diverse captions for clinical fetal ultrasound scan videos than those without gaze at the expense of the perceived sentence structure. The results also show that the generated captions are similar to sonographer speech in terms of discussing the visual content and the scanning actions performed.