Speech Processing for Language Learning: A Practical Approach to Computer-Assisted Pronunciation Teaching

Speech Processing for Language Learning: A Practical Approach to Computer-Assisted Pronunciation Teaching
复制标题

DOI:
10.3390/electronics10030235
复制
发表时间:
2021-02-01
期刊:
影响因子:
2.9
通讯作者:
Blake, John
Blake, John
中科院分区:
工程技术3区
文献类型:
--
作者:
Bogach, Natalia;Boitsova, Elena;Blake, John

文献摘要

被引文献

相似文献

本文讨论了当代计算机和信息技术如何帮助提高外语学习,不仅通过支持更好、更灵活的工作流程和数字化学习材料,还通过信号处理算法的技术改进创造出全新的用例。我们讨论了一种方法,并提出了一个整体的解决方案,以教学音素等对正确发音至关重要的语音现象;构成短语节奏的音节和停顿的能量和持续时间;以及话语中的语调运动,即短语语调。研究音调计算机辅助发音训练(CAPT)系统的工作原型是一种移动设备工具,它提供一组基于“听和重复”方法的任务,并实时给出视听反馈。本工作总结了为丰富当前版本的CAPT工具所做的努力,增加了两个新功能:语音转录和模式语音和学习者语音的节奏模式。两者都是基于第三方自动语音识别(ASR)库Kaldi设计的,该库被整合在studintonation信号处理软件核心中。我们还研究了自动语音识别在CAPT系统工作流程中的适用性范围,并评估了人类专家转录与我们代码中自动获得的转录之间的Levenstein距离。我们开发了一种基于声学和语言ASR模型的节奏重建算法。研究还表明,即使有足够正确的音素产生,学习者也不能产生正确的短语节奏和语调,因此,在单一的学习环境中对声音、节奏和语调进行联合训练是有益的。为了减轻录音缺陷,在处理的所有语音记录中应用了语音活动检测(VAD)。试验表明,studintonation可以创建转录和处理节奏模式,但在连接语音转录方面发现了一些具体问题。更新学习者在语音评价方面的反馈,将基于动态时间规整(DTW)的传统机制与交叉递归量化分析(CRQA)方法相结合,提高了学习者的识别能力。CRQA指标与DTW指标相结合,可以提高学习者绩效评估的准确性。讨论了计算机辅助英语语音教学的主要意义。
This article contributes to the discourse on how contemporary computer and information technology may help in improving foreign language learning not only by supporting better and more flexible workflow and digitizing study materials but also through creating completely new use cases made possible by technological improvements in signal processing algorithms. We discuss an approach and propose a holistic solution to teaching the phonological phenomena which are crucial for correct pronunciation, such as the phonemes; the energy and duration of syllables and pauses, which construct the phrasal rhythm; and the tone movement within an utterance, i.e., the phrasal intonation. The working prototype of StudyIntonation Computer-Assisted Pronunciation Training (CAPT) system is a tool for mobile devices, which offers a set of tasks based on a "listen and repeat" approach and gives the audio-visual feedback in real time. The present work summarizes the efforts taken to enrich the current version of this CAPT tool with two new functions: the phonetic transcription and rhythmic patterns of model and learner speech. Both are designed on a base of a third-party automatic speech recognition (ASR) library Kaldi, which was incorporated inside StudyIntonation signal processing software core. We also examine the scope of automatic speech recognition applicability within the CAPT system workflow and evaluate the Levenstein distance between the transcription made by human experts and that obtained automatically in our code. We developed an algorithm of rhythm reconstruction using acoustic and language ASR models. It is also shown that even having sufficiently correct production of phonemes, the learners do not produce a correct phrasal rhythm and intonation, and therefore, the joint training of sounds, rhythm and intonation within a single learning environment is beneficial. To mitigate the recording imperfections voice activity detection (VAD) is applied to all the speech records processed. The try-outs showed that StudyIntonation can create transcriptions and process rhythmic patterns, but some specific problems with connected speech transcription were detected. The learners feedback in the sense of pronunciation assessment was also updated and a conventional mechanism based on dynamic time warping (DTW) was combined with cross-recurrence quantification analysis (CRQA) approach, which resulted in a better discriminating ability. The CRQA metrics combined with those of DTW were shown to add to the accuracy of learner performance estimation. The major implications for computer-assisted English pronunciation teaching are discussed.