Leveraging Text Data for Word Segmentation for Underresourced Languages

Leveraging Text Data for Word Segmentation for Underresourced Languages
复制标题

DOI:
10.21437/interspeech.2017-1262
复制
发表时间:
2017-08
期刊:
--
影响因子:
--
通讯作者:
Thomas Glarner;Benedikt T. Boenninghoff;Oliver Walter;Reinhold Häb-Umbach
Thomas Glarner;Benedikt T. Boenninghoff;Oliver Walter;Reinhold Häb-Umbach
中科院分区:
其他
文献类型:
--
作者:
Thomas Glarner;Benedikt T. Boenninghoff;Oliver Walter;Reinhold Häb-Umbach

文献摘要

相似文献

在这篇文章中,我们展示了如何利用文本数据来支持在资源不足的目标语言中从音频输入中发现单词。给定一定数量的音频和附加的不相关的文本数据,该方法能够学习从声学单元到字符的概率映射,并利用它将音频数据分割成单词,而不需要发音词典。这是由三个组件实现的:无监督声学单元发现系统、受监督训练的声学单元到字素转换器、以及单词发现系统,该系统用在文本数据上训练的语言模型来初始化。多个设置的实验表明,使用文本数据初始化语言模型可以较大幅度地提高分词性能。
In this contribution we show how to exploit text data to support word discovery from audio input in an underresourced target language. Given audio, of which a certain amount is transcribed at the word level, and additional unrelated text data, the approach is able to learn a probabilistic mapping from acoustic units to characters and utilize it to segment the audio data into words without the need of a pronunciation dictionary. This is achieved by three components: an unsupervised acoustic unit discovery system, a supervisedly trained acoustic unit-to-grapheme converter, and a word discovery system, which is initialized with a language model trained on the text data. Experiments for multiple setups show that the initialization of the language model with text data improves the word segementation performance by a large margin.