Differentiable Allophone Graphs for Language-Universal Speech Recognition

Differentiable Allophone Graphs for Language-Universal Speech Recognition
复制标题

用于语言通用语音识别的可微分音素图

DOI:
--
复制
发表时间:
2021
期刊:
Interspeech
影响因子:
--
通讯作者:
Shinji Watanabe
Shinji Watanabe
中科院分区:
--
文献类型:
--
作者:
Brian Yan;Siddharth Dalmia;David R. Mortensen;Florian Metze;Shinji Watanabe

文献摘要

参考文献

被引文献

相似文献

建立语言通用的语音识别系统需要产生可以在不同语言之间共享的语音单元。虽然在语言特定的音素或表面级别的语音注释是容易获得的,但在通用音素级别的注释相对较少并且难以产生。在这项工作中,我们提出了一个通用的框架,以获得语音级的监督,只有音素transmittance和音素到音素的映射与可学习的权重表示使用加权有限状态转换器,我们称之为可微的allophone图。通过多语言训练,我们建立了一个通用的基于音素的语音识别模型,每个语言都具有可解释的概率音素到音素映射。语言学家可以使用这些基于音素的系统来记录新的语言,构建基于音素的词典,以捕捉丰富的发音变化,并重新评估所见语言的音素映射。我们证明了我们提出的框架的上述好处与7种不同的语言训练的系统。
Building language-universal speech recognition systems entails producing phonological units of spoken sound that can be shared across languages. While speech annotations at the language-specific phoneme or surface levels are readily available, annotations at a universal phone level are relatively rare and difficult to produce. In this work, we present a general framework to derive phone-level supervision from only phonemic transcriptions and phone-to-phoneme mappings with learnable weights represented using weighted finite-state transducers, which we call differentiable allophone graphs. By training multilingually, we build a universal phone-based speech recognition model with interpretable probabilistic phone-to-phoneme mappings for each language. These phone-based systems with learned allophone graphs can be used by linguists to document new languages, build phone-based lexicons that capture rich pronunciation variations, and re-evaluate the allophone mappings of seen language. We demonstrate the aforementioned benefits of our proposed framework with a system trained on 7 diverse languages.
AlloVera:多语言同位素数据库
DOI: --
发表时间: 2020
期刊: Proceedings of The 12th Language Resources and Evaluation Conference
影响因子: --
作者:
Mortensen, David R.;Li, Xinjian;Littell, Patrick;Michaud, Alexis;Rijhwani, Shruti;Anastasopoulos, Antonios;Black, Alan W;Metze, Florian;Neubig, Graham
通讯作者: Neubig, Graham
DOI: 10.1109/icassp40776.2020.9054362
发表时间: 2020-02
期刊: ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子: --
作者:
Xinjian Li;Siddharth Dalmia;Juncheng Billy Li;Matthew Russell Lee;Patrick Littell;Jiali Yao;Antonios Anastasopoulos;David R. Mortensen;Graham Neubig;A. Black;Florian Metze
通讯作者: Xinjian Li;Siddharth Dalmia;Juncheng Billy Li;Matthew Russell Lee;Patrick Littell;Jiali Yao;Antonios Anastasopoulos;David R. Mortensen;Graham Neubig;A. Black;Florian Metze