Align or attend? Toward More Efficient and Accurate Spoken Word Discovery Using Speech-to-Image Retrieval

Align or attend? Toward More Efficient and Accurate Spoken Word Discovery Using Speech-to-Image Retrieval
复制标题

DOI:
10.1109/icassp39728.2021.9414418
复制
发表时间:
2021-06
期刊:
ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Liming Wang;Xinsheng Wang;M. Hasegawa-Johnson;O. Scharenborg;N. Dehak
Liming Wang;Xinsheng Wang;M. Hasegawa-Johnson;O. Scharenborg;N. Dehak
中科院分区:
其他
文献类型:
--
作者:
Liming Wang;Xinsheng Wang;M. Hasegawa-Johnson;O. Scharenborg;N. Dehak

文献摘要

被引文献

相似文献

多模态词发现(MWD)通常被视为语音到图像检索问题的副产品。然而,我们的理论分析表明,某种对齐/注意机制是至关重要的MWD系统学习有意义的词级表示。我们通过在MSCOCO和Flickr 8 k上进行检索和单词发现实验来验证我们的理论,并从经验上证明,具有自我注意力的神经MT和统计MT的单词发现分数都上级最先进的神经检索系统,分别优于2%和5%的对齐F1分数。
Multimodal word discovery (MWD) is often treated as a byproduct of the speech-to-image retrieval problem. However, our theoretical analysis shows that some kind of alignment/attention mechanism is crucial for a MWD system to learn meaningful word-level representation. We verify our theory by conducting retrieval and word discovery experiments on MSCOCO and Flickr8k, and empirically demonstrate that both neural MT with self-attention and statistical MT achieve word discovery scores that are superior to those of a state-of-the-art neural retrieval system, outperforming it by 2% and 5% alignment F1 scores respectively.