Optimising Speaker-Dependent Feature Extraction Parameters to Improve Automatic Speech Recognition Performance for People with Dysarthria.

Optimising Speaker-Dependent Feature Extraction Parameters to Improve Automatic Speech Recognition Performance for People with Dysarthria.
复制标题

优化依赖说话者的特征提取参数,以改善构音障碍患者的自动语音识别性能。

DOI:
10.3390/s21196460
复制
发表时间:
2021-09-27
期刊:
Sensors (Basel, Switzerland)
影响因子:
--
通讯作者:
Fanucci L
Fanucci L
中科院分区:
其他
文献类型:
--
作者:
Marini M;Vanello N;Fanucci L

文献摘要

参考文献

被引文献

相似文献

在自动语音识别(ASR)系统中,面对言语障碍是一个很大的挑战,因为标准方法在存在构音障碍时是无效的。我们工作的第一个目的是确认一种新的语音分析技术对患有构音障碍的说话者的有效性。这种新方法利用了用于计算初始短时傅里叶变换的光谱分析窗口的大小和移位参数的微调,以提高依赖扬声器的ASR系统的性能。第二个目标是定义说话者的声音特征与最佳窗口和移位参数之间是否存在相关性,从而使ASR系统对特定说话者的误差最小化。在我们的实验中,我们使用了受损和未受损的意大利语。具体来说,我们使用了来自IDEA数据库的30名构音障碍患者和来自CLIPS数据库的10名专业演讲者。这两个数据库都是免费提供的。结果证实,如果标准的ASR系统对患有构音障碍的说话者表现不佳,则可以通过使用新的语音分析来改进。否则,新方法在未受损和低受损语言的情况下是无效的。此外,某些说话人的语音特征与其最优参数之间存在相关性。
Within the field of Automatic Speech Recognition (ASR) systems, facing impaired speech is a big challenge because standard approaches are ineffective in the presence of dysarthria. The first aim of our work is to confirm the effectiveness of a new speech analysis technique for speakers with dysarthria. This new approach exploits the fine-tuning of the size and shift parameters of the spectral analysis window used to compute the initial short-time Fourier transform, to improve the performance of a speaker-dependent ASR system. The second aim is to define if there exists a correlation among the speaker’s voice features and the optimal window and shift parameters that minimises the error of an ASR system, for that specific speaker. For our experiments, we used both impaired and unimpaired Italian speech. Specifically, we used 30 speakers with dysarthria from the IDEA database and 10 professional speakers from the CLIPS database. Both databases are freely available. The results confirm that, if a standard ASR system performs poorly with a speaker with dysarthria, it can be improved by using the new speech analysis. Otherwise, the new approach is ineffective in cases of unimpaired and low impaired speech. Furthermore, there exists a correlation between some speaker’s voice features and their optimal parameters.
DOI: 10.1016/j.jvoice.2004.02.005
发表时间: 2005-06-01
期刊: JOURNAL OF VOICE
影响因子: 2.2
作者:
Tanner, K;Roy, N;Buder, EH
通讯作者: Buder, EH
DOI: 10.1007/s10579-011-9145-0
发表时间: 2012-12-01
影响因子: 2.7
作者:
Rudzicz, Frank;Namasivayam, Aravind Kumar;Wolff, Talya
通讯作者: Wolff, Talya
DOI: 10.1561/2000000004
发表时间: 2007-01-01
影响因子: --
作者:
Gales, Mark;Young, Steve
通讯作者: Young, Steve
DOI: 10.1109/tnsre.2018.2802914
发表时间: 2018-03-01
影响因子: 4.9
作者:
Joy, Neethu Mariam;Umesh, S.
通讯作者: Umesh, S.
DOI: 10.1002/widm.2
发表时间: 2011-01-01
影响因子: 7.8
作者:
Rousseeuw, Peter J.;Hubert, Mia
通讯作者: Hubert, Mia