Subword unit representations for spoken document retrieval

Subword unit representations for spoken document retrieval
复制标题

用于语音文档检索的子词单元表示

DOI:
10.21437/eurospeech.1997-460
复制
发表时间:
1997
期刊:
--
影响因子:
--
通讯作者:
V. Zue
V. Zue
中科院分区:
--
文献类型:
--
作者:
Kenney Ng;V. Zue

文献摘要

被引文献

相似文献

本文研究了使用子词单元表示作为使用关键字识别或单词识别生成的单词的替代来进行口语文档检索的可行性。我们的研究是基于观察到的基于词的检索方法面临的问题,要么必须知道关键字来搜索先验,要么需要非常大的识别词汇量来覆盖不断增长和多样化的消息集合的内容。在这项研究中,我们考察了一系列来自语音转录的不同复杂程度的子词单位。基本的基本单位是音素;通过改变语音单位的细节程度和序列长度,可以得出更复杂和更不复杂的单位。我们测量DI(CID:11)个现有子词单元对E(CID:11)分别索引和检索大量记录的语音消息的能力。我们还比较了当潜在的语音转录是完美的和当它们包含语音识别错误时的性能。
This paper investigates the feasibility of using subword unit representations for spoken document retrieval as an alternative to using words generated by either keyword spotting or word recognition. Our investigation is motivated by the observation that word-based retrieval approaches face the problem of either having to know the keywords to search for a priori, or requiring a very large recognition vocabulary in order to cover the contents of growing and diverse message collections. In this study, we examine a range of subword units of varying complexity derived from phonetic transcriptions. The basic underlying unit is the phone; more and less complex units are derived by varying the level of detail and the length of sequences of the phonetic units. We measure the ability of the di(cid:11)erent subword units to e(cid:11)ectively index and retrieve a large collection of recorded speech messages. We also compare their performance when the underlying phonetic transcriptions are perfect and when they contain phonetic recognition errors.