Unsupervised Lexicon Acquisition from Speech and Text
Unsupervised Lexicon Acquisition from Speech and Text
复制标题
从语音和文本中获取无监督的词典
DOI:
10.1109/icassp.2007.366939
复制
发表时间:
2007
期刊:
影响因子:
--
通讯作者:
M. Nishimura
中科院分区:
文献类型:
--
作者:
Gakuto Kurata;Shinsuke Mori;N. Itoh;M. Nishimura
When introducing a large vocabulary continuous speech recognition (LVCSR) system into a specific domain, it is preferable to add the necessary domain-specific words and their correct pronunciations selectively to the lexicon, especially in the areas where the LVCSR system should be updated frequently by adding new words. In this paper, we propose an unsupervised method of word acquisition in Japanese, where no spaces exist between words. In our method, by taking advantage of the speech of the target domain, we selected the domain-specific words among an enormous number of word candidates extracted from the raw corpora. The experiments showed that the acquired lexicon was of good quality and that it contributed to the performance of the LVCSR system for the target domain.