A comparative study of adaptive, automatic recognition of disordered speech

A comparative study of adaptive, automatic recognition of disordered speech
复制标题

自适应自动识别紊乱语音的比较研究

DOI:
10.21437/interspeech.2012-484
复制
发表时间:
2012
期刊:
IEEE Transactions on Acoustics, Speech, and Signal Processing
影响因子:
--
通讯作者:
Thomas Hain
Thomas Hain
中科院分区:
--
文献类型:
--
作者:
H. Christensen;S. Cunningham;C. Fox;P. Green;Thomas Hain

文献摘要

被引文献

相似文献

语音驱动的辅助技术可以成为身体残疾人士传统界面的一种有吸引力的替代方案。然而,通常缺乏运动控制的语音发音导致紊乱的讲话,作为条件称为构音障碍。普通的自动语音识别产品通常不能为构音障碍者提供满意的语音识别效果,而构音障碍语音识别是一个日益活跃的研究领域。合适的数据稀疏是一个很大的挑战。这里描述的实验使用UA语音,最大的构音障碍数据库之一,这仍然是很容易的数量级小于典型的语音数据库。本研究调查了LVCSR社区开发的基本培训和适应技术可以带我们走多远。各种ASR系统使用最大似然和MAP适应策略建立与所有扬声器获得显着改善相比,基线系统,无论其条件的严重程度。最好的系统显示出平均34%的相对改善已知的公布结果。扬声器的可懂度和系统的类型,这将代表一个最佳的操作点的性能方面的相关性的分析表明,严重构音障碍的扬声器,系统配置的确切选择是更关键的扬声器与不那么无序的讲话。
Speech-driven assistive technology can be an attractive alternative to conventional interfaces for people with physical disabilities. However, often the lack of motor-control of the speech articulators results in disordered speech, as condition known as dysarthria. Dysarthric speakers can generally not obtain satisfactory performances with off-the-shelf automatic speech recognition (ASR) products and disordered speech ASR is an increasingly active research area. Sparseness of suitable data is a big challenge. The experiments described here use UAspeech, one of the largest dysarthric databases available, which is still easily an order of magnitude smaller than typical speech databases. This study investigates how far fundamental training and adaptation techniques developed in the LVCSR community can take us. A variety of ASR systems using maximum likelihood and MAP adaptation strategies are established with all speakers obtaining significant improvements compared to the baseline system regardless of the severity of their condition. The best systems show on average 34% relative improvement on known published results. An analysis of the correlation between intelligibility of the speaker and the type of system which would represent an optimal operating point in terms of performance shows that for severely dysarthric speakers, the exact choice of system configuration is more critical than for speakers with less disordered speech.