Assessment of Non-Native Speech Intelligibility using Wav2vec2-based Mispronunciation Detection and Multi-level Goodness of Pronunciation Transformer

Assessment of Non-Native Speech Intelligibility using Wav2vec2-based Mispronunciation Detection and Multi-level Goodness of Pronunciation Transformer
复制标题

DOI:
10.21437/interspeech.2023-2371
复制
发表时间:
2023-08
期刊:
--
影响因子:
--
通讯作者:
R. Shekar;Mu Yang;K. Hirschi;Stephen Looney;Okim Kang;J. Hansen
R. Shekar;Mu Yang;K. Hirschi;Stephen Looney;Okim Kang;J. Hansen
中科院分区:
其他
文献类型:
--
作者:
R. Shekar;Mu Yang;K. Hirschi;Stephen Looney;Okim Kang;J. Hansen

文献摘要

相似文献

在计算机辅助语音训练(CAPT)中,语音自动评价(APA)在为自主语言学习者提供反馈方面发挥着重要作用。一些基于端到端音素识别的错误发音检测和诊断系统已经取得了很好的效果。然而,评估第二语言(L2)的可理解性仍然是一个具有挑战性的问题。一个问题是缺乏来自非母语人士的大规模标记语音数据。此外,仅依靠语音层面的一个方面(如准确性)可能无法充分评估语音质量和二语可理解性。可以利用音段/语音级别的特征,如发音良好度(GOP),然而,特征粒度可能会导致韵律级别(超音段)发音评估的差异。本研究采用基于Wav2vec 2.0的MDD和基于发音良好度特征的Transformer来表征二语可读性。在这里,一个带有人工标注的韵律(超分段)标签的第二语言语音数据集被用于多粒度和多方面的语音评估和识别第二语言英语语音可理解性的重要因素。该研究对自动发音分数与超分音特征和听者感知之间的关系进行了变革性的比较评估,这些综合起来有助于为二语学习者提供即时评估工具和解决方案的开发。
Automatic pronunciation assessment (APA) plays an important role in providing feedback for self-directed language learners in computer-assisted pronunciation training (CAPT). Several mispronunciation detection and diagnosis (MDD) systems have achieved promising performance based on end-to-end phoneme recognition. However, assessing the intelligibility of second language (L2) remains a challenging problem. One issue is the lack of large-scale labeled speech data from non-native speakers. Additionally, relying only on one aspect (e.g., accuracy) at a phonetic level may not provide a sufficient assessment of pronunciation quality and L2 intelligibility. It is possible to leverage segmental/phonetic-level features such as goodness of pronunciation (GOP), however, feature granularity may cause a discrepancy in prosodic-level (suprasegmental) pronunciation assessment. In this study, Wav2vec 2.0-based MDD and Good-ness Of Pronunciation feature-based Transformer are employed to characterize L2 intelligibility. Here, an L2 speech dataset, with human-annotated prosodic (suprasegmental) labels, is used for multi-granular and multi-aspect pronunciation assessment and identification of factors important for intelligibility in L2 English speech. The study provides a transformative comparative assessment of automated pronunciation scores versus the relationship between suprasegmental features and listener perceptions, which taken collectively can help support the development of instantaneous assessment tools and solutions for L2 learners.