Leveraging Text Data for Word Segmentation for Underresourced Languages
Leveraging Text Data for Word Segmentation for Underresourced Languages
复制标题
DOI:
10.21437/interspeech.2017-1262
复制
发表时间:
2017-08
期刊:
影响因子:
--
通讯作者:
Thomas Glarner;Benedikt T. Boenninghoff;Oliver Walter;Reinhold Häb-Umbach
中科院分区:
文献类型:
--
作者:
Thomas Glarner;Benedikt T. Boenninghoff;Oliver Walter;Reinhold Häb-Umbach
In this contribution we show how to exploit text data to support word discovery from audio input in an underresourced target language. Given audio, of which a certain amount is transcribed at the word level, and additional unrelated text data, the approach is able to learn a probabilistic mapping from acoustic units to characters and utilize it to segment the audio data into words without the need of a pronunciation dictionary. This is achieved by three components: an unsupervised acoustic unit discovery system, a supervisedly trained acoustic unit-to-grapheme converter, and a word discovery system, which is initialized with a language model trained on the text data. Experiments for multiple setups show that the initialization of the language model with text data improves the word segementation performance by a large margin.