A Dual-Attention Learning Network With Word and Sentence Embedding for Medical Visual Question Answering

A Dual-Attention Learning Network With Word and Sentence Embedding for Medical Visual Question Answering
复制标题

DOI:
10.1109/tmi.2023.3322868
复制
发表时间:
2024-02-01
影响因子:
10.6
通讯作者:
Gong,Hongfang
Gong,Hongfang
中科院分区:
工程技术1区
文献类型:
--
作者:
Huang,Xiaofei;Gong,Hongfang

文献摘要

相似文献

医学视觉问答系统的研究有助于计算机辅助诊断的发展。MVQA是一项旨在基于给定的医学图像和相关的自然语言问题来预测准确且令人信服的答案的任务。这一任务要求提取医学知识丰富的特征内容,并对它们进行细粒度理解。因此,构造有效的特征提取和理解方案是建模的关键。现有的MVQA问题抽取方案主要关注单词信息,忽略了文本中的医学信息,如医学概念和领域术语。同时,一些视觉和文本特征理解方案不能有效地捕捉区域和关键字之间的相关性,以进行合理的视觉推理。提出了一种基于单词和句子嵌入的双注意学习网络(DALNet-WSE)。我们设计了一个模块,Transformer with Sentence Embedding(TSE),以提取包含关键字和医疗信息的问题的双嵌入表示。提出了一种由自我注意和引导注意组成的双注意学习模型,用于建模密集的通道内和通道间交互。通过多DAL模块(DALs),学习视觉和文本的共同注意可以提高理解的粒度,改善视觉推理。在ImageCLEF 2019 VQA-MED(VQA-MED 2019)和VQA-RAD数据集上的实验结果表明,我们提出的方法优于以前的最先进方法。根据消融研究和Grad-CAM图,DALNet-WSE能够提取出丰富的文本信息,具有较强的视觉推理能力。
Research in medical visual question answering (MVQA) can contribute to the development of computer-aided diagnosis. MVQA is a task that aims to predict accurate and convincing answers based on given medical images and associated natural language questions. This task requires extracting medical knowledge-rich feature content and making fine-grained understandings of them. Therefore, constructing an effective feature extraction and understanding scheme are keys to modeling. Existing MVQA question extraction schemes mainly focus on word information, ignoring medical information in the text, such as medical concepts and domain-specific terms. Meanwhile, some visual and textual feature understanding schemes cannot effectively capture the correlation between regions and keywords for reasonable visual reasoning. In this study, a dual-attention learning network with word and sentence embedding (DALNet-WSE) is proposed. We design a module, transformer with sentence embedding (TSE), to extract a double embedding representation of questions containing keywords and medical information. A dual-attention learning (DAL) module consisting of self-attention and guided attention is proposed to model intensive intramodal and intermodal interactions. With multiple DAL modules (DALs), learning visual and textual co-attention can increase the granularity of understanding and improve visual reasoning. Experimental results on the ImageCLEF 2019 VQA-MED (VQA-MED 2019) and VQA-RAD datasets demonstrate that our proposed method outperforms previous state-of-the-art methods. According to the ablation studies and Grad-CAM maps, DALNet-WSE can extract rich textual information and has strong visual reasoning ability.