Prediction of Creaky Speech by Recurrent Neural Networks Using Psychoacoustic Roughness

Prediction of Creaky Speech by Recurrent Neural Networks Using Psychoacoustic Roughness
复制标题

DOI:
10.1109/jstsp.2019.2949422
复制
发表时间:
2020-02
影响因子:
7.5
通讯作者:
J. Villegas;K. Markov;Jeremy Perkins;Seunghun J. Lee
J. Villegas;K. Markov;Jeremy Perkins;Seunghun J. Lee
中科院分区:
工程技术1区
文献类型:
--
作者:
J. Villegas;K. Markov;Jeremy Perkins;Seunghun J. Lee

文献摘要

被引文献

相似文献

据报道,使用心理声学粗糙度模型作为嘎吱声的预测器。我们发现声音片段的粗糙度时间分布可以预测语音中是否存在嘎吱声。使用简单的双向递归神经网络 (rnn),我们能够仅根据粗糙度轨迹来预测声音片段中是否存在嘎吱声,其准确度与使用至少 12 维输入数据(包括前两个谐波之间的幅度差异、残余峰值突出度等)训练的 rnn 获得的准确度相似。结合粗糙度和多维输入数据训练 rnn 提高了预测器的性能,但并不显着。同样,通过输入特征的时间导数来扩充数据集并不能提高预测器的性能。所提出的基于粗糙度的预测器简化了语料库中嘎吱声的解释和比较,并表明粗糙度预测模型可以成功地用于语音中嘎吱声间隔的分类。
The use of a psychoacoustic roughness model as a predictor of creaky voice is reported. We found that the roughness temporal profile of vocalic segments can predict the presence of creakiness in speech. Using a simple bi-directional Recurrent Neural Network (rnn), we were able to predict the presence of creakiness in vocalic segments from only roughness traces with an accuracy similar to that obtained with rnns trained on at least 12-dimensional input data (including amplitude difference between the first two harmonics, residual peak prominence, etc.). Training rnns with the combination of roughness and multidimensional input data improved the performance of the predictor, but not significantly. Likewise, augmenting the dataset by time derivatives of the input features did not improve the predictor's performance. The proposed roughness-based predictor eases interpretation and comparison of creakiness among corpora and suggests that roughness prediction models could be successfully used for classification of creaky intervals in speech.