Near-Optimal Active Learning for Multilingual Grapheme-to-Phoneme Conversion

Near-Optimal Active Learning for Multilingual Grapheme-to-Phoneme Conversion
复制标题

多语言字素到音素转换的近乎最优主动学习

DOI:
--
复制
发表时间:
2023
期刊:
影响因子:
--
通讯作者:
Licheng Wu
Licheng Wu
中科院分区:
--
文献类型:
--
作者:
Dezhi Cao;Yue Zhao;Licheng Wu

文献摘要

参考文献

相似文献

语音词典的构建依赖于高质量和广泛的数据驱动的训练数据。然而,用于此目的的语料库的手动标注既昂贵又耗时,特别是对于缺乏足够数据和资源的低资源语言。多语种语音词典中包含了一些常见的音素或语音单元,这意味着这些音素或单元在不同语言的发音中具有相似性,可以用于低资源语言的语音词典的构建过程。通过使用多语言发音词典,可以在不同语言之间共享知识,从而提高低资源语言发音词典的质量和准确性。在本文中,我们建议使用共享的发音特征之间的多种语言,构建一个通用的音素集,然后使用它来标记的话,为多种语言。为了实现这一目标,我们首先开发了一个基于编码器-解码器深度神经网络的字素-音素(G2 P)模型。然后,我们在建立发音词典的过程中采用了一种近似最优的主动学习方法,从一个大的,未标记的语料库中选择有信息的样本,并由专家标记。实验表明,该方法选择了约1/5的未标记数据,并取得了比大数据训练方法更高的转换精度。通过选择性地标记模型中具有高不确定性的样本,同时避免标记由当前模型准确预测的样本,我们的方法大大提高了发音词典构建的效率。
The construction of pronunciation dictionaries relies on high-quality and extensive training data in data-driven way. However, the manual annotation of corpus for this purpose is both costly and time consuming, especially for low-resource languages that lack sufficient data and resources. A multilingual pronunciation dictionary includes some common phonemes or phonetic units, which means that these phonemes or units have similarities in the pronunciation of different languages and can be used in the construction process of pronunciation dictionaries for low-resource languages. By using a multilingual pronunciation dictionary, knowledge can be shared among different languages, thus improving the quality and accuracy of pronunciation dictionaries for low-resource languages. In this paper, we propose using shared articulatory features among multiple languages to construct a universal phoneme set, which is then used to label words for multiple languages. To achieve this, we first developed a grapheme−phoneme (G2P) model based on an encoder−decoder deep neural network. We then adopted a near-optimal active learning method in the process of building the pronunciation dictionary to select informative samples from a large, unlabeled corpus and had them labeled by experts. Our experiments demonstrate that this method selected about 1/5 of the unlabeled data and achieved an even higher conversion accuracy than the results of the large data training method. By selectively labeling samples with a high uncertainty in the model, while avoiding labeling samples that were accurately predicted by the current model, our method greatly enhances the efficiency of pronunciation dictionary construction.
DOI: 10.1136/amiajnl-2011-000648
发表时间: 2012-09-01
影响因子: 6.4
作者:
Figueroa, Rosa L.;Zeng-Treitler, Qing;Wiechmann, Eduardo P.
通讯作者: Wiechmann, Eduardo P.