Representations of language in a model of visually grounded speech signal

Representations of language in a model of visually grounded speech signal
复制标题

DOI:
10.18653/v1/p17-1057
复制
发表时间:
2017-02
期刊:
--
影响因子:
--
通讯作者:
Grzegorz Chrupała;Lieke Gelderloos;A. Alishahi
Grzegorz Chrupała;Lieke Gelderloos;A. Alishahi
中科院分区:
其他
文献类型:
--
作者:
Grzegorz Chrupała;Lieke Gelderloos;A. Alishahi

文献摘要

被引文献

相似文献

我们提出了一个视觉接地模型的语音感知项目口语和图像的联合语义空间。我们使用多层递归高速公路网络来模拟口语的时间性质,并表明它可以从输入信号中提取基于形式和意义的语言知识。我们对训练模型的不同组件所使用的表示进行了深入分析,并表明语义方面的编码往往会随着层次的上升而变得更加丰富,而语言输入的形式相关方面的编码往往会最初增加,然后趋于平稳或减少。
We present a visually grounded model of speech perception which projects spoken utterances and images to a joint semantic space. We use a multi-layer recurrent highway network to model the temporal nature of spoken speech, and show that it learns to extract both form and meaning-based linguistic knowledge from the input signal. We carry out an in-depth analysis of the representations used by different components of the trained model and show that encoding of semantic aspects tends to become richer as we go up the hierarchy of layers, whereas encoding of form-related aspects of the language input tends to initially increase and then plateau or decrease.