Towards Visually Grounded Sub-word Speech Unit Discovery

Towards Visually Grounded Sub-word Speech Unit Discovery
复制标题

走向基于视觉的子词语音单元发现

DOI:
--
复制
发表时间:
2019
期刊:
IEEE International Conference on Acoustics, Speech, and Signal Processing
影响因子:
--
通讯作者:
James R. Glass
James R. Glass
中科院分区:
--
文献类型:
--
作者:
David F. Harwath;James R. Glass

文献摘要

被引文献

相似文献

在本文中,我们研究了在卷积神经网络模型中出现可解释的子词语音单元的方式,该模型被训练来将原始语音波形与语义相关的自然图像场景相关联。我们展示了如何从模型中间层的激活模式中表面地提取双音素边界,这表明该模型可能正在利用这些事件来实现单词识别的目的。我们提出了一系列实验,研究由这些事件编码的信息。
In this paper, we investigate the manner in which interpretable sub-word speech units emerge within a convolutional neural network model trained to associate raw speech waveforms with semantically related natural image scenes. We show how diphone boundaries can be superficially extracted from the activation patterns of intermediate layers of the model, suggesting that the model may be leveraging these events for the purpose of word recognition. We present a series of experiments investigating the information encoded by these events.