Video-guided Machine Translation with Spatial Hierarchical Attention Network

Video-guided Machine Translation with Spatial Hierarchical Attention Network
复制标题

DOI:
10.18653/v1/2021.acl-srw.9
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
Weiqi Gu;Haiyue Song;Chenhui Chu;S. Kurohashi
Weiqi Gu;Haiyue Song;Chenhui Chu;S. Kurohashi
中科院分区:
其他
文献类型:
--
作者:
Weiqi Gu;Haiyue Song;Chenhui Chu;S. Kurohashi

文献摘要

相似文献

视频引导机器翻译作为多模态机器翻译的一种,其目的是利用视频内容作为辅助信息来解决机器翻译中的词义歧义问题。以往的研究仅使用预训练动作检测模型的特征作为视频的运动表征来解决动词意义歧义,而名词意义歧义则是一个问题。为了解决这个问题,我们提出了一个视频引导机器翻译系统,通过在视频中使用空间和运动表示。对于空间特征,我们提出了一个分层关注网络来对从对象级到视频级的空间信息进行建模。在VATEX数据集上的实验表明,我们的系统达到了35.86的BLEU-4分数,比SOTA方法的单一模型高0.51分。
Video-guided machine translation, as one type of multimodal machine translations, aims to engage video contents as auxiliary information to address the word sense ambiguity problem in machine translation. Previous studies only use features from pretrained action detection models as motion representations of the video to solve the verb sense ambiguity, leaving the noun sense ambiguity a problem. To address this problem, we propose a video-guided machine translation system by using both spatial and motion representations in videos. For spatial features, we propose a hierarchical attention network to model the spatial information from object-level to video-level. Experiments on the VATEX dataset show that our system achieves 35.86 BLEU-4 score, which is 0.51 score higher than the single model of the SOTA method.