Korean Dialect Identification Based on Intonation Modeling

Korean Dialect Identification Based on Intonation Modeling
复制标题

基于语调建模的朝鲜语方言识别

DOI:
--
复制
发表时间:
2021
期刊:
Oriental COCOSDA International Conference on Speech Database and Assessments
影响因子:
--
通讯作者:
Minhwa Chung
Minhwa Chung
中科院分区:
--
文献类型:
--
作者:
Jooyoung Lee;K. Kim;Minhwa Chung

文献摘要

被引文献

相似文献

韩国方言识别(K-DID)是一项具有挑战性的任务,由于其相对未开发的研究领域,方言之间的相互理解,以及缺乏足够的韩国方言数据集在过去。随着大规模的方言数据集,本文提出了语调建模的韩国方言喂养帧的声学特征的神经网络的顺序建模。与以前基于音节的音高标记的韵律标记相比,我们的语调建模方法是通过组合一组频谱特征来实现的,包括基频,在具有注意力机制的双向LSTM网络上训练。我们相信,注意机制,使检测方言丰富的部分隐藏在占主导地位的非方言段在同一话语。我们测试了不同组合的扬声器年龄和讲话风格的网络。K-DID的最佳性能是在话语级准确率达到68.51%,这超过了我们以前的工作。
Korean dialect identification (K-DID) is a challenging task due to its relatively unexplored field of study, mutual comprehensibility between the dialects, and lack of sufficient Korean dialect datasets available in the past. With large-scaled dialect datasets now available, this paper proposes intonational modeling of the Korean dialects by feeding frame-wise acoustic features on sequential modeling of a neural network. Compared to previous prosodic labeling with syllable-based pitch marking, our approach of intonation modeling is realized with the combination of a set of spectral features, including fundamental frequency, trained on a bidirectional LSTM network with attention mechanism. We believe the attention mechanism enables the detection of dialect-rich segments hidden among the dominant non-dialect segments within the same utterance. We test the networks on different combinations of speaker ages and speech styles. The best performance of the K-DID is achieved with 68.51 % in utterance-level accuracy, which surpasses our previous work.