Automatic Prediction of Intelligibility of Words and Phonemes Produced Orally by Japanese Learners of English

Automatic Prediction of Intelligibility of Words and Phonemes Produced Orally by Japanese Learners of English
复制标题

DOI:
10.1109/slt54892.2023.10023307
复制
发表时间:
2023-01
期刊:
2022 IEEE Spoken Language Technology Workshop (SLT)
影响因子:
--
通讯作者:
Chuanbo Zhu-;Takuya Kunihara;D. Saito;N. Minematsu;Noriko Nakanishi
Chuanbo Zhu-;Takuya Kunihara;D. Saito;N. Minematsu;Noriko Nakanishi
中科院分区:
其他
文献类型:
--
作者:
Chuanbo Zhu-;Takuya Kunihara;D. Saito;N. Minematsu;Noriko Nakanishi

文献摘要

相似文献

语言学习的实际目标是与他人顺利沟通,许多教师非常注重测量可理解性而不是口音,通常被视为实际理解的正确性。然而,可理解性的自动预测还没有得到很好的发展,特别是对于较小的单位,如单词和音素。这主要是因为难以测量听者的边听行为,因此很难建立足够规模的二语语料库,并进行可理解性标注来训练基于网络的预测器。在本文中,我们使用小延迟的口头听写(即阴影)来注释可理解性,以从两个不同语言背景的评分者中收集足够大的语料库。由于感知的可理解性取决于他们的语言背景,因此应该考虑到他们之间的差异。因此,利用该语料库,建立了一个多评分者神经模型来预测二语语音中每个评分者对单个单词和音素的可理解性。研究了两项任务,即可理解性分数的回归和给定片段的可理解性或不可理解性分类。结果表明,我们的模型的F1得分高于内部评分协议,表明我们的模型可以很好地模拟两个评分者,尽管他们具有不同的语言背景。
The practical goal for language learning is smooth communication with others, and many teachers have a strong focus on measurement of not accentedness but intelligibility, often regarded as correctness of actual understanding. However, automatic prediction of intelligibility has not been well developed especially for smaller units such as words and phonemes. This is mainly because of difficulty of measuring while-listening behaviors of listeners, and thus it was difficult to build an L2 speech corpus of a sufficient size with intelligibility annotation to train a network-based predictor. In this paper, we annotate intelligibility using oral dictation with a small delay, i.e., shadowing, to collect a large enough corpus from two raters with different language backgrounds. Since perceived intelligibility depends on their language background, inter-rater difference should be taken into account. Therefore with this corpus, a multi-rater neural model is built to predict each rater's intelligibility of the individual words and phonemes in L2 speech. Two tasks are examined, i.e., regression of intelligibility scores and classification of a given segment to be intelligible or not. Results show that our model has higher F1 scores than intra-rater agreements, indicating that our model can simulate the two raters accurately well although they have different language background.