Semantic Query-by-example Speech Search Using Visual Grounding

Semantic Query-by-example Speech Search Using Visual Grounding
复制标题

DOI:
10.1109/icassp.2019.8683275
复制
发表时间:
2019-04
期刊:
ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
H. Kamper;Aristotelis Anastassiou;Karen Livescu
H. Kamper;Aristotelis Anastassiou;Karen Livescu
中科院分区:
其他
文献类型:
--
作者:
H. Kamper;Aristotelis Anastassiou;Karen Livescu

文献摘要

被引文献

相似文献

最近的一些研究已经开始研究如何通过在训练时利用伴随的图像来训练语音系统。任务的示例包括关键字预测和模式内和跨模式检索。在这里,我们考虑这样的模型可以用于查询的例子(QBE)搜索,检索相关的一个给定的口头查询的话语的任务。我们特别感兴趣的语义QBE,其中的任务不仅是检索话语包含确切的查询实例,但也话语的意义是相关的查询。我们遵循分段QBE方法,其中可变持续时间的语音段(查询,搜索话语)被映射到固定维度的嵌入向量。我们表明,使用嵌入函数训练视觉接地语音数据的QBE系统优于一个纯粹的声学QBE系统的准确和语义检索性能。
A number of recent studies have started to investigate how speech systems can be trained on untranscribed speech by leveraging accompanying images at training time. Examples of tasks include keyword prediction and within- and across-mode retrieval. Here we consider how such models can be used for query-by-example (QbE) search, the task of retrieving utterances relevant to a given spoken query. We are particularly interested in semantic QbE, where the task is not only to retrieve utterances containing exact instances of the query, but also utterances whose meaning is relevant to the query. We follow a segmental QbE approach where variable-duration speech segments (queries, search utterances) are mapped to fixed-dimensional embedding vectors. We show that a QbE system using an embedding function trained on visually grounded speech data outperforms a purely acoustic QbE system in terms of both exact and semantic retrieval performance.