Epitran: Precision G2P for Many Languages

Epitran: Precision G2P for Many Languages
复制标题

DOI:
--
复制
发表时间:
2018-05
影响因子:
2.5
通讯作者:
David R. Mortensen;Siddharth Dalmia;Patrick Littell
David R. Mortensen;Siddharth Dalmia;Patrick Littell
中科院分区:
计算机科学3区
文献类型:
--
作者:
David R. Mortensen;Siddharth Dalmia;Patrick Littell

文献摘要

被引文献

相似文献

Epitran 是一个用于 G2P(字素到音素)转换的大规模多语言、多后端系统,支持 61 种语言。它采用语言正字法中的单词标记并输出 IPA 或 X-SAMPA 中的音素表示。主系统是用Python编写的,并且作为开源软件公开提供。其功效已在多个与语言迁移、多语言模型和语音相关的研究项目中得到证明。在特定的 ASR 任务中,Epitran 被证明可以提高声学建模的 Babel 基线的单词错误率。
Epitran is a massively multilingual, multiple back-end system for G2P (grapheme-to-phoneme) transduction which is distributed with support for 61 languages. It takes word tokens in the orthography of a language and outputs a phonemic representation in either IPA or X-SAMPA. The main system is written in Python and is publicly available as open source software. Its efficacy has been demonstrated in multiple research projects relating to language transfer, polyglot models, and speech. In a particular ASR task, Epitran was shown to improve the word error rate over Babel baselines for acoustic modeling.