Unsupervised Lexicon Acquisition from Speech and Text

Unsupervised Lexicon Acquisition from Speech and Text
复制标题

从语音和文本中获取无监督的词典

DOI:
10.1109/icassp.2007.366939
复制
发表时间:
2007
期刊:
2007 IEEE International Conference on Acoustics, Speech and Signal Processing - ICASSP '07
影响因子:
--
通讯作者:
M. Nishimura
M. Nishimura
中科院分区:
--
文献类型:
--
作者:
Gakuto Kurata;Shinsuke Mori;N. Itoh;M. Nishimura

文献摘要

被引文献

相似文献

在将大词汇量连续语音识别(LVCSR)系统引入特定领域时,最好选择性地将必要的特定于领域的单词及其正确发音添加到词典中,特别是在LVCSR系统需要经常通过添加新词来更新的领域。在本文中,我们提出了一种无监督的日语单词习得方法,其中单词之间不存在空格。在我们的方法中,我们利用目标领域的语音,从原始语料库中提取的大量候选词中选择特定于该领域的词。实验表明,所获得的词典质量良好,有助于LVCSR系统在目标领域的性能。
When introducing a large vocabulary continuous speech recognition (LVCSR) system into a specific domain, it is preferable to add the necessary domain-specific words and their correct pronunciations selectively to the lexicon, especially in the areas where the LVCSR system should be updated frequently by adding new words. In this paper, we propose an unsupervised method of word acquisition in Japanese, where no spaces exist between words. In our method, by taking advantage of the speech of the target domain, we selected the domain-specific words among an enormous number of word candidates extracted from the raw corpora. The experiments showed that the acquired lexicon was of good quality and that it contributed to the performance of the LVCSR system for the target domain.