FluentNet: End-to-End Detection of Stuttered Speech Disfluencies With Deep Learning
FluentNet: End-to-End Detection of Stuttered Speech Disfluencies With Deep Learning
复制标题
DOI:
10.1109/taslp.2021.3110146
复制
发表时间:
2021-01-01
影响因子:
5.4
通讯作者:
Etemad, Ali
中科院分区:
文献类型:
--
作者:
Kourkounakis, Tedd;Hajavi, Amirhossein;Etemad, Ali
Millions of people are affected by stuttering and other speech disfluencies, with the majority of the world having experienced mild stutters while communicating under stressful conditions. While there has been much research in the field of automatic speech recognition and language models, stutter detection and recognition has not received as much attention. To this end, we propose an end-to-end deep neural network, FluentNet, capable of detecting a number of different stutter types. FluentNet consists of a Squeeze-and-Excitation Residual convolutional neural network which facilitate the learning of strong spectral frame-level representations, followed by a set of bidirectional long short-term memory layers that aid in learning effective temporal relationships. Lastly, FluentNet uses an attention mechanism to focus on the important parts of speech to obtain a better performance. We perform a number of different experiments, comparisons, and ablation studies to evaluate our model. Our model achieves state-of-the-art results by outperforming other solutions in the field on the publicly available UCLASS dataset. Additionally, we present LibriStutter: a stuttered speech dataset based on the public LibriSpeech dataset with synthesized stutters. We also evaluate FluentNet on this dataset, showing the strong performance of our model versus a number of baseline and state-of-the-art techniques.