Machine learning techniques for semantic analysis of dysarthric speech: An experimental study

Machine learning techniques for semantic analysis of dysarthric speech: An experimental study
复制标题

DOI:
10.1016/j.specom.2018.04.005
复制
发表时间:
2018-05-01
影响因子:
3.2
通讯作者:
Haeb-Umbach, Reinhold
Haeb-Umbach, Reinhold
中科院分区:
计算机科学3区
文献类型:
--
作者:
Despotovic, Vladimir;Walter, Oliver;Haeb-Umbach, Reinhold

文献摘要

被引文献

相似文献

我们提出了一个实验比较七个国家的最先进的机器学习算法的任务的语义分析的口语输入,特别强调的应用程序构音障碍的语音。构音障碍是一种运动性言语障碍,其特征是音素的清晰度差。为了满足这些非规范音素实现,我们采用了一种无监督学习方法来估计语音识别的声学模型,它不需要对训练数据进行文字转录。即使对于随后的语义分析任务,也仅采用弱监督,由此训练话语仅伴随有语义标签,而不是字面转录。两个数据库,其中一个包含构音障碍的语音,结果表明,马尔可夫逻辑网络和条件随机场大大优于其他机器学习方法。马尔可夫逻辑网络已被证明是特别强大的识别错误,这是由不精确的发音障碍的语音。
We present an experimental comparison of seven state-of-the-art machine learning algorithms for the task of semantic analysis of spoken input, with a special emphasis on applications for dysarthric speech. Dysarthria is a motor speech disorder, which is characterized by poor articulation of phonemes. In order to cater for these non-canonical phoneme realizations, we employed an unsupervised learning approach to estimate the acoustic models for speech recognition, which does not require a literal transcription of the training data. Even for the subsequent task of semantic analysis, only weak supervision is employed, whereby the training utterance is accompanied by a semantic label only, rather than a literal transcription. Results on two databases, one of them containing dysarthric speech, are presented showing that Markov logic networks and conditional random fields substantially outperform other machine learning approaches. Markov logic networks have proved to be especially robust to recognition errors, which are caused by imprecise articulation in dysarthric speech.