Towards Visually Grounded Sub-word Speech Unit Discovery
Towards Visually Grounded Sub-word Speech Unit Discovery
复制标题
走向基于视觉的子词语音单元发现
DOI:
--
复制
发表时间:
2019
期刊:
影响因子:
--
通讯作者:
James R. Glass
中科院分区:
文献类型:
--
作者:
David F. Harwath;James R. Glass
In this paper, we investigate the manner in which interpretable sub-word speech units emerge within a convolutional neural network model trained to associate raw speech waveforms with semantically related natural image scenes. We show how diphone boundaries can be superficially extracted from the activation patterns of intermediate layers of the model, suggesting that the model may be leveraging these events for the purpose of word recognition. We present a series of experiments investigating the information encoded by these events.