Cascade Attention Guided Residue Learning GAN for Cross-Modal Translation

Cascade Attention Guided Residue Learning GAN for Cross-Modal Translation
复制标题

DOI:
10.1109/icpr48806.2021.9412890
复制
发表时间:
2019-07
期刊:
2020 25th International Conference on Pattern Recognition (ICPR)
影响因子:
--
通讯作者:
Bin Duan;Wei Wang;Hao Tang;Hugo Latapie;Yan Yan-Yan
Bin Duan;Wei Wang;Hao Tang;Hugo Latapie;Yan Yan-Yan
中科院分区:
其他
文献类型:
--
作者:
Bin Duan;Wei Wang;Hao Tang;Hugo Latapie;Yan Yan-Yan

文献摘要

被引文献

相似文献

当我们还是婴儿的时候,我们就直觉地发展出了将来自不同认知传感器的输入(如视觉、音频和文本)关联起来的能力。然而,在机器学习中,这种跨模态学习是一项重要的任务,因为不同的模态没有同质的属性。以往的工作发现,应该有不同的模态之间的桥梁。从神经学和心理学的角度来看,人类有能力将一种模态与另一种模态联系起来,例如,把一只鸟的图片和它的歌声联系起来,反之亦然。机器学习算法是否有可能在给定音频信号的情况下恢复场景?在本文中,我们提出了一种新的级联注意力引导残差GAN(CAR-GAN),旨在重建相应的音频信号的场景。特别地,我们提出了一个残差模块来逐步缩小不同模态之间的差距。此外,一个级联的注意力引导网络与一个新的分类损失函数的设计,以解决跨模态学习任务。我们的模型保持了高层语义标签域的一致性,并能够平衡两种不同的模态。实验结果表明,我们的模型在具有挑战性的Sub-URMP数据集上实现了最先进的跨模态视听生成。
Since we were babies, we intuitively develop the ability to correlate the input from different cognitive sensors such as vision, audio, and text. However, in machine learning, this cross-modal learning is a nontrivial task because different modalities have no homogeneous properties. Previous works discover that there should be bridges among different modalities. From a neurology and psychology perspective, humans have the capacity to link one modality with another one, e.g., associating a picture of a bird with the only hearing of its singing and vice versa. Is it possible for machine learning algorithms to recover the scene given the audio signal? In this paper, we propose a novel Cascade Attention-Guided Residue GAN (CAR-GAN), aiming at reconstructing the scenes given the corresponding audio signals. Particularly, we present a residue module to mitigate the gap between different modalities progressively. Moreover, a cascade attention guided network with a novel classification loss function is designed to tackle the cross-modal learning task. Our model keeps consistency in the high-level semantic label domain and is able to balance two different modalities. The experimental results demonstrate that our model achieves the state-of-the-art cross-modal audio-visual generation on the challenging Sub-URMP dataset.