Recurrent Networks for Guided Multi-Attention Classification

Recurrent Networks for Guided Multi-Attention Classification
复制标题

DOI:
10.1145/3394486.3403083
复制
发表时间:
2020-07
期刊:
Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining
影响因子:
--
通讯作者:
Xin Dai;Xiangnan Kong;Tian Guo;J. B. Lee;Xinyue Liu;C. Moore
Xin Dai;Xiangnan Kong;Tian Guo;J. B. Lee;Xinyue Liu;C. Moore
中科院分区:
其他
文献类型:
--
作者:
Xin Dai;Xiangnan Kong;Tian Guo;J. B. Lee;Xinyue Liu;C. Moore

文献摘要

相似文献

近年来,基于注意力的图像分类越来越受到人们的重视。用于基于注意力的分类的现有方法通常需要大的训练集,并且在假设图像的标签仅依赖于图像中的单个对象(即感兴趣区域)的情况下操作。然而,在许多实际应用中(例如医学成像),收集大量训练集是非常昂贵的。此外,每幅图像的标签通常由多个感兴趣区域(ROI)共同确定。幸运的是,对于这样的应用,通常可以收集每个训练图像中的ROI的位置。本文研究了导引多注意分类问题,其目标是在(1)小样本和(2)每幅图像具有多个感兴趣区的双重约束下获得高精度。我们提出了一种用于多注意分类的模型,称为引导注意递归网络(GARN)。与现有的基于注意力的方法不同,GARN利用关于多个ROI的指导信息,从而使其即使在小样本大小的情况下也能很好地工作。对三种不同视觉任务的实验研究表明,我们的引导注意方法可以有效地提高多注意图像分类的模型性能。
Attention-based image classification has gained increasing popularity in recent years. State-of-the-art methods for attention-based classification typically require a large training set and operate under the assumption that the label of an image depends solely on a single object (i.e. region of interest) in the image. However, in many real-world applications (e.g. medical imaging), it is very expensive to collect a large training set. Moreover, the label of each image is usually determined jointly by multiple regions of interest (ROIs). Fortunately, for such applications, it is often possible to collect the locations of the ROIs in each training image. In this paper, we study the problem of guided multi-attention classification, the goal of which is to achieve high accuracy under the dual constraints of (1) small sample size, and (2) multiple ROIs for each image. We propose a model, called Guided Attention Recurrent Network (GARN), for multi-attention classification. Different from existing attention-based methods, GARN utilizes guidance information regarding multiple ROIs thus allowing it to work well even when sample size is small. Empirical studies on three different visual tasks show that our guided attention approach can effectively boost model performance for multi-attention image classification.