Bidirectional Relationship Inferring Network for Referring Image Localization and Segmentation

Bidirectional Relationship Inferring Network for Referring Image Localization and Segmentation
复制标题

DOI:
10.1109/tnnls.2021.3106153
复制
发表时间:
2021-09
影响因子:
10.4
通讯作者:
Guang Feng;Zhiwei Hu;Lihe Zhang;Jiayu Sun;Huchuan Lu
Guang Feng;Zhiwei Hu;Lihe Zhang;Jiayu Sun;Huchuan Lu
中科院分区:
计算机科学1区
文献类型:
--
作者:
Guang Feng;Zhiwei Hu;Lihe Zhang;Jiayu Sun;Huchuan Lu

文献摘要

相似文献

近年来,参考图像的定位与分割引起了人们的广泛关注。然而,现有的方法缺乏对语言和视觉之间相互依存关系的清晰描述。为此,我们提出了一个双向关系推理网络(BRINet)来有效地解决具有挑战性的任务。具体而言,我们首先使用视觉引导的语言注意模块来感知每个图像区域对应的关键词。然后,语言引导视觉注意采用习得的自适应语言来引导视觉特征的更新。它们共同构成双向跨模态注意模块(BCAM),实现语言与视觉的相互引导。它们可以帮助网络更好地对齐跨模式特征。在传统语言引导视觉注意的基础上,我们进一步设计了一种非对称语言引导视觉注意,通过对每个像素和每个池子区域之间的关系进行建模,显著降低了计算成本。此外,利用分割导向的自底向上增强模块(SBAM),有选择地组合多级信息流进行目标定位。实验表明,该方法在3个参考图像定位数据集和4个参考图像分割数据集上优于其他最先进的方法。
Recently, referring image localization and segmentation has aroused widespread interest. However, the existing methods lack a clear description of the interdependence between language and vision. To this end, we present a bidirectional relationship inferring network (BRINet) to effectively address the challenging tasks. Specifically, we first employ a vision-guided linguistic attention module to perceive the keywords corresponding to each image region. Then, language-guided visual attention adopts the learned adaptive language to guide the update of the visual features. Together, they form a bidirectional cross-modal attention module (BCAM) to achieve the mutual guidance between language and vision. They can help the network align the cross-modal features better. Based on the vanilla language-guided visual attention, we further design an asymmetric language-guided visual attention, which significantly reduces the computational cost by modeling the relationship between each pixel and each pooled subregion. In addition, a segmentation-guided bottom-up augmentation module (SBAM) is utilized to selectively combine multilevel information flow for object localization. Experiments show that our method outperforms other state-of-the-art methods on three referring image localization datasets and four referring image segmentation datasets.