Learning Neural Audio Embeddings for Grounding Semantics in Auditory Perception

Learning Neural Audio Embeddings for Grounding Semantics in Auditory Perception
复制标题

DOI:
10.1613/jair.5665
复制
发表时间:
2017-12
期刊:
J. Artif. Intell. Res.
影响因子:
--
通讯作者:
Douwe Kiela;S. Clark
Douwe Kiela;S. Clark
中科院分区:
其他
文献类型:
--
作者:
Douwe Kiela;S. Clark

文献摘要

相似文献

多模态语义旨在在感知中奠定语义表示的基础,它依赖于特征规范或原始图像数据来进行感知输入。在本文中,我们使用多模态语义的标准评估来研究原始听觉数据中的基础语义表示。在展示了这种基于听觉的表征的质量之后,我们展示了如何将它们应用于与听觉感知相关的任务,包括两个无监督的分类实验,并提供进一步的分析。我们发现从深度神经网络转移的特征优于音频词袋方法。据我们所知,这是第一个结合深度神经网络提取的文本信息和听觉信息构建多模态模型的工作,也是第一个评估三模态(文本、视觉和听觉)语义模型性能的工作。
Multi-modal semantics, which aims to ground semantic representations in perception, has relied on feature norms or raw image data for perceptual input. In this paper we examine grounding semantic representations in raw auditory data, using standard evaluations for multi-modal semantics. After having shown the quality of such auditorily grounded representations, we show how they can be applied to tasks where auditory perception is relevant, including two unsupervised categorization experiments, and provide further analysis. We find that features transfered from deep neural networks outperform bag of audio words approaches. To our knowledge, this is the first work to construct multi-modal models from a combination of textual information and auditory information extracted from deep neural networks, and the first work to evaluate the performance of tri-modal (textual, visual and auditory) semantic models.