FluentNet: End-to-End Detection of Stuttered Speech Disfluencies With Deep Learning

FluentNet: End-to-End Detection of Stuttered Speech Disfluencies With Deep Learning
复制标题

DOI:
10.1109/taslp.2021.3110146
复制
发表时间:
2021-01-01
影响因子:
5.4
通讯作者:
Etemad, Ali
Etemad, Ali
中科院分区:
计算机科学2区
文献类型:
--
作者:
Kourkounakis, Tedd;Hajavi, Amirhossein;Etemad, Ali

文献摘要

被引文献

相似文献

数以百万计的人受到口吃和其他言语不流利的影响,世界上大多数人在紧张的条件下交流时都经历过轻微的口吃。虽然在自动语音识别和语言模型领域已经有了很多研究,但口吃检测和识别还没有得到那么多的关注。为此,我们提出了一种端到端的深度神经网络,FluentNet,能够检测许多不同类型的口吃。FluentNet由一个压缩和激励残差卷积神经网络组成,它促进了强光谱帧级别表示的学习,随后是一组有助于学习有效时间关系的双向长期短期记忆层。最后,FluentNet使用一种注意力机制将注意力集中在重要的词性上,以获得更好的性能。我们进行了许多不同的实验、比较和消融研究来评估我们的模型。我们的模型在公开可用的UCLASS数据集上的表现优于该领域的其他解决方案,从而获得了最先进的结果。此外,我们还提出了LibriStutter:一个基于带有合成卡顿的公共LibriSpeech数据集的卡顿语音数据集。我们还在此数据集上评估了FluentNet,显示了我们的模型相对于一些基准和最先进技术的强大性能。
Millions of people are affected by stuttering and other speech disfluencies, with the majority of the world having experienced mild stutters while communicating under stressful conditions. While there has been much research in the field of automatic speech recognition and language models, stutter detection and recognition has not received as much attention. To this end, we propose an end-to-end deep neural network, FluentNet, capable of detecting a number of different stutter types. FluentNet consists of a Squeeze-and-Excitation Residual convolutional neural network which facilitate the learning of strong spectral frame-level representations, followed by a set of bidirectional long short-term memory layers that aid in learning effective temporal relationships. Lastly, FluentNet uses an attention mechanism to focus on the important parts of speech to obtain a better performance. We perform a number of different experiments, comparisons, and ablation studies to evaluate our model. Our model achieves state-of-the-art results by outperforming other solutions in the field on the publicly available UCLASS dataset. Additionally, we present LibriStutter: a stuttered speech dataset based on the public LibriSpeech dataset with synthesized stutters. We also evaluate FluentNet on this dataset, showing the strong performance of our model versus a number of baseline and state-of-the-art techniques.