Multi-lingual Transformer Training for Khmer Automatic Speech Recognition
Multi-lingual Transformer Training for Khmer Automatic Speech Recognition
复制标题
DOI:
10.1109/apsipaasc47483.2019.9023137
复制
发表时间:
2019-11
期刊:
影响因子:
--
通讯作者:
Kak Soky;Sheng Li;Tatsuya Kawahara;Sopheap Seng
中科院分区:
文献类型:
--
作者:
Kak Soky;Sheng Li;Tatsuya Kawahara;Sopheap Seng
Currently, there are three challenges for constructing reliable ASR systems for the Khmer language: (1) the lack of language resources (text and speech corpora) in digital form, (2) the writing system without explicit word boundary, and (3) the pronunciation model is not well studied. In this paper, to avoid the extensive work on selecting proper acoustic units (e.g., phones, syllables) and preparing the frame-level labels on the traditional DNN-HMM framework, we directly use words or characters as the label using state-of-the-art transformer-based end-to-end model. Moreover, we use the multi-lingual training framework to tackle the low-resource data problem. All experiments are performed on the Basic Expressions Travel Corpus (BTEC) datasets. The experiments show that the proposed multi-lingual transformer-based end-to-end model can achieve significant improvement compared to the DNN-HMM baseline model11The work was performed during Mr. Kak Soky was in NIPTICT. He is currently with Ministry of Education, Youth, and Sports (MoEYS), Cambodia.