Multi-lingual Transformer Training for Khmer Automatic Speech Recognition

Multi-lingual Transformer Training for Khmer Automatic Speech Recognition
复制标题

DOI:
10.1109/apsipaasc47483.2019.9023137
复制
发表时间:
2019-11
期刊:
2019 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC)
影响因子:
--
通讯作者:
Kak Soky;Sheng Li;Tatsuya Kawahara;Sopheap Seng
Kak Soky;Sheng Li;Tatsuya Kawahara;Sopheap Seng
中科院分区:
其他
文献类型:
--
作者:
Kak Soky;Sheng Li;Tatsuya Kawahara;Sopheap Seng

文献摘要

相似文献

目前,构建可靠的高棉语 ASR 系统面临三个挑战:(1)缺乏数字形式的语言资源(文本和语音语料库);(2)没有明确词边界的书写系统;(3)发音模型没有得到很好的研究。在本文中,为了避免在传统 DNN-HMM 框架上选择合适的声学单元(例如音素、音节)和准备帧级标签的大量工作,我们使用最先进的基于 Transformer 的端到端模型直接使用单词或字符作为标签。此外,我们使用多语言训练框架来解决资源不足的数据问题。所有实验均在基本表达旅游语料库 (BTEC) 数据集上进行。实验表明,所提出的基于多语言 Transformer 的端到端模型与 DNN-HMM 基线模型相比可以取得显着的改进11这项工作是 Kak Soky 先生在 NIPTICT 期间进行的。他目前供职于柬埔寨教育、青年和体育部 (MoEYS)。
Currently, there are three challenges for constructing reliable ASR systems for the Khmer language: (1) the lack of language resources (text and speech corpora) in digital form, (2) the writing system without explicit word boundary, and (3) the pronunciation model is not well studied. In this paper, to avoid the extensive work on selecting proper acoustic units (e.g., phones, syllables) and preparing the frame-level labels on the traditional DNN-HMM framework, we directly use words or characters as the label using state-of-the-art transformer-based end-to-end model. Moreover, we use the multi-lingual training framework to tackle the low-resource data problem. All experiments are performed on the Basic Expressions Travel Corpus (BTEC) datasets. The experiments show that the proposed multi-lingual transformer-based end-to-end model can achieve significant improvement compared to the DNN-HMM baseline model11The work was performed during Mr. Kak Soky was in NIPTICT. He is currently with Ministry of Education, Youth, and Sports (MoEYS), Cambodia.