Q2ATransformer: Improving Medical VQA via an Answer Querying Decoder

Q2ATransformer: Improving Medical VQA via an Answer Querying Decoder
复制标题

Q2ATransformer:通过答案查询解码器改进医疗 VQA

DOI:
--
复制
发表时间:
2023
期刊:
Information Processing in Medical Imaging
影响因子:
--
通讯作者:
Luping Zhou
Luping Zhou
中科院分区:
--
文献类型:
--
作者:
Yunyi Liu;Zhanyu Wang;Dong Xu;Luping Zhou

文献摘要

参考文献

被引文献

相似文献

医学视觉问答系统(VQA)对理解医学图像所承载的临床相关信息起着辅助作用。针对医学图像的问题包括两类:封闭式(如是/否问题)和开放式。为了获得答案,大多数现有的医学VQA方法依赖于分类方法,而少数作品试图使用生成方法或两者的混合。分类方法相对简单,但在长开放式问题上表现不佳。为了弥补这一差距,在本文中,我们提出了一个新的Transformer为基础的框架医学VQA(命名为Q2 A Transformer),它集成了分类和生成方法的优点,并提供了一个统一的处理封闭端和开放端的问题。具体来说,我们引入了一个额外的Transformer解码器与一组可学习的候选答案嵌入查询存在的每个答案类到一个给定的图像问题对。通过Transformer注意力,候选答案嵌入与图像-问题对的融合特征相互作用以做出决策。通过这种方式,尽管是基于分类的方法,但我们的方法提供了一种与答案信息进行交互的机制,以便像基于生成的方法那样进行预测。另一方面,通过分类,我们通过减少答案的搜索空间来降低任务难度。我们的方法在两个医学VQA基准上实现了新的最先进的性能。特别是对于开放式问题,我们在VQA-RAD和PathVQA上分别达到了79.19%和54.85%,分别有16.09%和41.45%的绝对改进。
Medical Visual Question Answering (VQA) systems play a supporting role to understand clinic-relevant information carried by medical images. The questions to a medical image include two categories: close-end (such as Yes/No question) and open-end. To obtain answers, the majority of the existing medical VQA methods relies on classification approaches, while a few works attempt to use generation approaches or a mixture of the two. The classification approaches are relatively simple but perform poorly on long open-end questions. To bridge this gap, in this paper, we propose a new Transformer based framework for medical VQA (named as Q2ATransformer), which integrates the advantages of both the classification and the generation approaches and provides a unified treatment for the close-end and open-end questions. Specifically, we introduce an additional Transformer decoder with a set of learnable candidate answer embeddings to query the existence of each answer class to a given image-question pair. Through the Transformer attention, the candidate answer embeddings interact with the fused features of the image-question pair to make the decision. In this way, despite being a classification-based approach, our method provides a mechanism to interact with the answer information for prediction like the generation-based approaches. On the other hand, by classification, we mitigate the task difficulty by reducing the search space of answers. Our method achieves new state-of-the-art performance on two medical VQA benchmarks. Especially, for the open-end questions, we achieve 79.19% on VQA-RAD and 54.85% on PathVQA, with 16.09% and 41.45% absolute improvements, respectively.
DOI: --
发表时间: 2020-03
期刊: ArXiv
影响因子: --
作者:
Xuehai He;Yichen Zhang;Luntian Mou;E. Xing;P. Xie
通讯作者: Xuehai He;Yichen Zhang;Luntian Mou;E. Xing;P. Xie