Improving Mispronunciation Detection with Wav2vec2-based Momentum Pseudo-Labeling for Accentedness and Intelligibility Assessment

Improving Mispronunciation Detection with Wav2vec2-based Momentum Pseudo-Labeling for Accentedness and Intelligibility Assessment
复制标题

DOI:
10.48550/arxiv.2203.15937
复制
发表时间:
2022-03
期刊:
--
影响因子:
--
通讯作者:
Mu Yang;K. Hirschi;S. Looney;Okim Kang;J. Hansen
Mu Yang;K. Hirschi;S. Looney;Okim Kang;J. Hansen
中科院分区:
其他
文献类型:
--
作者:
Mu Yang;K. Hirschi;S. Looney;Okim Kang;J. Hansen

文献摘要

被引文献

相似文献

当前领先的发音错误检测和诊断(MDD)系统通过端到端音素识别实现了令人鼓舞的性能。这种端到端解决方案面临的挑战之一是自然 L2 语音中缺乏人工注释的音素。在这项工作中,我们通过伪标记 (PL) 过程利用未标记的 L2 语音,并扩展基于预训练的自监督学习 (SSL) 模型的微调方法。具体来说,我们使用 Wav2vec 2.0 作为我们的 SSL 模型,并使用原始标记的 L2 语音样本加上创建的伪标记 L2 语音样本对其进行微调。我们的伪标签是动态的,由在线模型的集合即时生成,这确保了我们的模型对伪标签噪声具有鲁棒性。我们表明,与仅标记样本的微调基线相比,使用伪标签进行微调可降低 5.35% 的音素错误率,并将 MDD F1 分数提高 2.48%。所提出的 PL 方法也被证明优于传统的离线 PL 方法。与最先进的 MDD 系统相比,我们的 MDD 解决方案可产生更准确、一致的语音错误诊断。此外,我们对单独的 UTD-4Accents 数据集进行了开放测试,其中我们的系统识别输出基于口音和清晰度,显示出与人类感知的强相关性。
Current leading mispronunciation detection and diagnosis (MDD) systems achieve promising performance via end-to-end phoneme recognition. One challenge of such end-to-end solutions is the scarcity of human-annotated phonemes on natural L2 speech. In this work, we leverage unlabeled L2 speech via a pseudo-labeling (PL) procedure and extend the fine-tuning approach based on pre-trained self-supervised learning (SSL) models. Specifically, we use Wav2vec 2.0 as our SSL model, and fine-tune it using original labeled L2 speech samples plus the created pseudo-labeled L2 speech samples. Our pseudo labels are dynamic and are produced by an ensemble of the online model on-the-fly, which ensures that our model is robust to pseudo label noise. We show that fine-tuning with pseudo labels achieves a 5.35% phoneme error rate reduction and 2.48% MDD F1 score improvement over a labeled-samples-only fine-tuning baseline. The proposed PL method is also shown to outperform conventional offline PL methods. Compared to the state-of-the-art MDD systems, our MDD solution produces a more accurate and consistent phonetic error diagnosis. In addition, we conduct an open test on a separate UTD-4Accents dataset, where our system recognition outputs show a strong correlation with human perception, based on accentedness and intelligibility.