Audiovisual Distinctive-Feature-Based Recognition of Dysarthric Speech
Audiovisual Distinctive-Feature-Based Recognition of Dysarthric Speech
批准号:
0534106
负责人:
Mark Hasegawa-Johnson
金额:
$0.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2005
资助国家:
美国
项目状态:
已结题
起止时间:
2005-11-15 至 2009-10-31
中文摘要
具有相当高的单词识别精度的自动听写软件现在已广泛提供给公众。然而,许多有大运动障碍的人,包括一些脑瘫和闭合性头部损伤的人,并没有享受到这些进步的好处,因为他们的一般运动障碍包括音障碍的一个组成部分,也就是说,由神经运动障碍引起的语言清晰度降低,而运动障碍经常妨碍正常使用键盘。由于这个原因,诵读困难的用户现在经常发现使用小词汇量的自动语音识别系统更容易,该系统使用表示字母和格式化命令的码字,以及精心适应个人用户语音的声学语音识别模型。但是,这种个性化语音识别系统的开发仍然是非常耗费人力的,因为人们对语言障碍的一般特征知之甚少。在本项目中,PI将研究诵读困难语音中发音错误的一般视听特征,并将研究结果应用于独立于说话人的大词汇和小词汇的诵读困难语音识别系统的开发。更具体地说,PI将研究基于单词的、基于电话的和基于语音特征的音频和视听语音识别模型,这些模型适用于小词汇和大词汇的语音识别器,专为个人计算机上的无限制文本输入而设计。这些模型将基于语音平衡的语音样本的音频和视频分析,这些样本来自一组患有构音障碍的说话者,分为以下四组:非常低的可理解性(0-25%的可理解性,由人类听众评分),低的可理解性(25-50%),中等的可理解性(50-75%)和高的可理解性(75-100%)。交互式语音分析将试图描述构音障碍中发音错误的说话者依赖特征;基于对初步数据的分析,PI假设发音错误的方式、发音错误的地点和发音错误是大约独立的事件。初步实验还表明,不同的构音障碍使用者将需要截然不同的语音识别架构,因为构音障碍的症状因人而异,因此PI将为构音障碍使用者开发和测试至少三类纯音频和视听语音识别算法:使用隐马尔可夫模型(hmm)的基于电话和全词识别器,使用支持向量机(svm)的语音特征和全词识别器,以及支持向量机- hmm混合识别器。我们将对这些模型进行评估,以确定每种算法的总体识别准确率、学习导致的准确率变化、构音障碍严重程度导致的准确率组间差异,以及准确率对词汇量的依赖。更广泛的影响:本研究将为构建具有神经运动障碍的计算机用户实际使用的语音识别工具奠定基础。在这个项目中开发的工具和数据都将是开源的,并且将被设计成可以很容易地移植到一个开源的视听语音识别系统中,为困难的用户服务。这项工作也可能具有目标社区以外的适用性,因为项目结果可能与许多其他在培训当前ASR系统方面有困难的人群(例如,有外国口音的人)相关。
英文摘要
Automatic dictation software with reasonably high word recognition accuracy is now widely available to the general public. Many people with gross motor impairments, including some people with cerebral palsy and closed head injuries, have not enjoyed the benefit of these advances, however, because their general motor impairment includes a component of dysarthria, that is to say reduced speech intelligibility caused by neuro-motor impairment, while the motor impairment often precludes normal use of a keyboard. For this reason, dysarthric users often now find it easier to use a small-vocabulary automatic speech recognition system, with code words representing letters and formatting commands, and with acoustic speech recognition models carefully adapted to the speech of the individual user. But development of such individualized speech recognition systems remains extremely labor-intensive, because so little is understood about the general characteristics of dysarthric speech. In this project, the PI will study the general audio and visual characteristics of articulation errors in dysarthric speech, and apply the results to the development of speaker-independent large-vocabulary and small-vocabulary audio and audiovisual dysarthric speech recognition systems. More specifically, the PI will research word-based, phone-based, and phonologic-feature-based audio and audiovisual speech recognition models for both small-vocabulary and large-vocabulary speech recognizers designed for unrestricted text entry on a personal computer. The models will be based on audio and video analysis of phonetically balanced speech samples from a group of speakers with dysarthria, categorized into the following four groups: very low intelligibility (0-25% intelligibility, as rated by human listeners), low intelligibility (25-50%), moderate intelligibility (50-75%), and high intelligibility (75-100%). Interactive phonetic analysis will seek to describe the talker-dependent characteristics of articulation error in dysarthria; based on analysis of preliminary data, the PI hypothesizes that manner of articulation errors, place of articulation errors, and voicing errors are approximately independent events. Preliminary experiments also suggest that different dysarthric users will require dramatically different speech recognition architectures, because the symptoms of dysarthria vary so much from subject to subject, so the PI will develop and test at least three categories of audio-only and audiovisual speech recognition algorithms for dysarthric users: phone-based and whole-word recognizers using hidden Markov models (HMMs), phonologic-feature-based and whole-word recognizers using support vector machines (SVMs), and hybrid SVM-HMM recognizers. The models will be evaluated to determine overall recognition accuracy of each algorithm, changes in accuracy due to learning, group differences in accuracy due to severity of dysarthria, and dependence of accuracy on vocabulary size.Broader Impacts: This research will lay the foundation for constructing a speech recognition tool for practical use by computer users with neuro-motor disabilities. Tools and data developed in this project will all be released open-source, and will be designed so they can be easily ported to an open-source audiovisual speech recognition system for dysarthric users. The work may also have applicability beyond the target community, in that project outcomes may be relevant to many other populations (e.g., people with foreign accents) who have trouble training current ASR systems.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
FAI: A New Paradigm for the Evaluation and Training of Inclusive Automatic Speech Recognition
-
批准号:2147350
-
项目类别:Standard Grant
-
资助金额:$50.0万
-
财政年份:2022
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
RI: Small: Collaborative Research: Automatic Creation of New Speech Sound Inventories
-
批准号:1910319
-
项目类别:Standard Grant
-
资助金额:$25.98万
-
财政年份:2019
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
EAGER: Matching Non-Native Transcribers to the Distinctive Features of the Language Transcribed
-
批准号:1550145
-
项目类别:Standard Grant
-
资助金额:$15.0万
-
财政年份:2015
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
FODAVA-Partner: Visualizing Audio for Anomaly Detection
-
批准号:0807329
-
项目类别:Continuing Grant
-
资助金额:$45.0万
-
财政年份:2008
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
RI Medium: Audio Diarization - Towards Comprehensive Description of Audio Events
-
批准号:0803219
-
项目类别:Standard Grant
-
资助金额:$24.99万
-
财政年份:2008
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
Prosodic, Intonational, and Voice Quality Correlates of Disfluency
-
批准号:0414117
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2004
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
CAREER: Landmark-Based Speech Recognition in Music and Speech Backgrounds
-
批准号:0132900
-
项目类别:Continuing Grant
-
资助金额:$39.58万
-
财政年份:2002
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
海外基金