Context-Aware Domain Adaptation in Semantic Segmentation

Context-Aware Domain Adaptation in Semantic Segmentation
复制标题

DOI:
10.1109/wacv48630.2021.00056
复制
发表时间:
2020-03
期刊:
2021 IEEE Winter Conference on Applications of Computer Vision (WACV)
影响因子:
--
通讯作者:
Jinyu Yang;Weizhi An;Chao-chao Yan;P. Zhao;Junzhou Huang
Jinyu Yang;Weizhi An;Chao-chao Yan;P. Zhao;Junzhou Huang
中科院分区:
其他
文献类型:
--
作者:
Jinyu Yang;Weizhi An;Chao-chao Yan;P. Zhao;Junzhou Huang

文献摘要

被引文献

相似文献

在本文中,我们考虑语义分割中的无监督域适应问题。该领域存在两个主要问题,即跨两个域转移域知识的内容和方式。现有方法主要集中于通过对抗学习(如何转移)来适应域不变特征(转移什么)。上下文依赖性对语义分割至关重要,然而,其可转移性仍未得到很好的理解。此外,如何跨两个域转移上下文信息仍未被探索。基于此,我们提出一种基于自注意力的交叉注意力机制,以捕捉两个域之间的上下文依赖性并适应可转移的上下文。为实现这一目标,我们设计了两个跨域注意力模块,从空间和通道两个视角来适应上下文依赖性。具体而言,空间注意力模块捕捉源图像和目标图像中每个位置之间的局部特征依赖性。通道注意力模块对每对跨域通道图之间的语义依赖性进行建模。为了适应上下文依赖性,我们进一步有选择性地聚合来自两个域的上下文信息。我们的方法相对于现有最先进方法的优越性在“GTA5到Cityscapes”以及“SYNTHIA到Cityscapes”上得到了实验验证。
In this paper, we consider the problem of unsupervised domain adaptation in the semantic segmentation. There are two primary issues in this field, i.e., what and how to transfer domain knowledge across two domains. Existing methods mainly focus on adapting domain-invariant features (what to transfer) through adversarial learning (how to transfer). Context dependency is essential for semantic segmentation, however, its transferability is still not well understood. Furthermore, how to transfer contextual information across two domains remains unexplored. Motivated by this, we propose a cross-attention mechanism based on self-attention to capture context dependencies between two domains and adapt transferable context. To achieve this goal, we design two cross-domain attention modules to adapt context dependencies from both spatial and channel views. Specifically, the spatial attention module captures local feature dependencies between each position in the source and target image. The channel attention module models semantic dependencies between each pair of cross-domain channel maps. To adapt context dependencies, we further selectively aggregate the context information from two domains. The superiority of our method over existing state-of-the-art methods is empirically proved on "GTA5 to Cityscapes" and "SYNTHIA to Cityscapes".