Universal Phone Recognition with a Multilingual Allophone System

Universal Phone Recognition with a Multilingual Allophone System
复制标题

DOI:
10.1109/icassp40776.2020.9054362
复制
发表时间:
2020-02
期刊:
ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Xinjian Li;Siddharth Dalmia;Juncheng Billy Li;Matthew Russell Lee;Patrick Littell;Jiali Yao;Antonios Anastasopoulos;David R. Mortensen;Graham Neubig;A. Black;Florian Metze
Xinjian Li;Siddharth Dalmia;Juncheng Billy Li;Matthew Russell Lee;Patrick Littell;Jiali Yao;Antonios Anastasopoulos;David R. Mortensen;Graham Neubig;A. Black;Florian Metze
中科院分区:
其他
文献类型:
--
作者:
Xinjian Li;Siddharth Dalmia;Juncheng Billy Li;Matthew Russell Lee;Patrick Littell;Jiali Yao;Antonios Anastasopoulos;David R. Mortensen;Graham Neubig;A. Black;Florian Metze

文献摘要

被引文献

相似文献

多语言模型可以通过跨语言共享参数来改进语言处理,特别是在资源匮乏的情况下。然而,多语言声学模型通常忽略了音素(在特定语言中支持词汇对比的声音)与其对应的音素(实际说出的声音,与语言无关)之间的差异。当组合各种训练语言时,这可能导致性能下降,因为相同注释的音素实际上可能对应于几种不同的底层语音实现。在这项工作中,我们提出了一个独立于语言的电话和依赖于语言的音素分布的联合模型。在11种语言的多语言ASR实验中,我们发现该模型在低资源条件下将音素错误率绝对提高了2%。此外,由于我们正在明确地建模与语言无关的手机,我们可以构建一个(几乎)通用的手机识别器,当与PHOIBLE[1]大型手动管理的手机库存数据库相结合时,可以定制成2000个与语言相关的识别器。在两种资源匮乏的本土语言Inuktitut和Tusom上的实验表明,我们的识别器实现了超过17%的电话准确率提高,向世界上所有语言的语音识别又迈进了一步
Multilingual models can improve language processing, particularly for low resource situations, by sharing parameters across languages. Multilingual acoustic models, however, generally ignore the difference between phonemes (sounds that can support lexical contrasts in a particular language) and their corresponding phones (the sounds that are actually spoken, which are language independent). This can lead to performance degradation when combining a variety of training languages, as identically annotated phonemes can actually correspond to several different underlying phonetic realizations. In this work, we propose a joint model of both language-independent phone and language-dependent phoneme distributions. In multilingual ASR experiments over 11 languages, we find that this model improves testing performance by 2% phoneme error rate absolute in low-resource conditions. Additionally, because we are explicitly modeling language-independent phones, we can build a (nearly-)universal phone recognizer that, when combined with the PHOIBLE [1] large, manually curated database of phone inventories, can be customized into 2,000 language dependent recognizers. Experiments on two low-resourced indigenous languages, Inuktitut and Tusom, show that our recognizer achieves phone accuracy improvements of more than 17%, moving a step closer to speech recognition for all languages in the world.1