课题基金 / 基金详情

UNS: Collaborative Research: Prosodic Control of Speech Synthesis for Assistive Communication in Severe Paralysis

UNS: Collaborative Research: Prosodic Control of Speech Synthesis for Assistive Communication in Severe Paralysis
UNS:合作研究:语音合成的韵律控制用于严重瘫痪的辅助沟通
批准号:
1509791
负责人:
Susan Koch Fager
金额:
$11.2万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-07-15 至 2018-06-30

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
1510563(Stepp)1509791(Koch Fager)这项工作将开发和评估一个系统,允许由于严重瘫痪而无法理解语音的个人控制语音合成器,包括韵律(音调,响度和持续时间的变化,传达意义)。这种对合成语音的改进以及用户对其控制的容易性将促进临床通信系统的改进功能,从而提高用户的生活质量。自然和可理解的语言生产在这些人将增加他们的能力,积极参与社会,并授权他们自我主张自己的医疗管理。本研究的目的是测试的假设,提供用户的替代和增强通信(AAC)使用韵律控制的方法将导致语音合成对收听者更自然并且为用户提供更大的功能。高达1.2%的人口由于中风或其他神经损伤而无法使用典型的语言来满足日常沟通需求,需要AAC来满足他们的沟通需求。他们的生活质量在很大程度上取决于能否获得这种通信,无论是为了社会互动还是传递有关紧急医疗需求的信息。最先进的AAC设备包含语音合成,允许用户与他人进行口头交流。然而,由此产生的合成语音既不自然又难以让其他人理解,并且通常被描述为“机器人”。具体来说,合成语音在音高、响度或节奏方面没有变化,而这些韵律特征在典型语音中用于传递情绪状态、话语形式(陈述与问题)、讽刺和强调。要求AAC用户单独控制这些维度中的每一个将导致难以忍受的缓慢和复杂的系统,对于已经大大降低了通信速率的个人来说,这是不可接受的负担。相反,这个项目将利用这样一个事实,即典型的语音可预测地使用这些韵律标记(音高,响度,节奏)。将开发一种新的AAC接口,允许用户修改合成语音输出的整体“压力”作为一个单一的维度,以提供易于控制,自然,和可理解的语音合成。联合PI将利用他们在语音技术、AAC临床应用和人机界面实时控制方面的综合专业知识,推动AAC技术的重大进步,以实现三个目标。在研究目标1中,将通过一种新的交互过程创建用于拼接语音合成的多重音语音库,在该过程中,健康说话者的语音生成被“误解”,从而促使说话者在重复响应中自然地强调特定的目标声音。这将产生一组三音子(基于周围声音,具有特定左右上下文的声音),其中包含所有可能的声音和重音组合。研究目标2是开发一个AAC界面,允许用户使用二维光标控制(例如,头部跟踪、眼睛跟踪),其中各个音素的重音将基于光标停留时间。在研究目标3中,AAC界面的功能将通过测试其对沟通交互自然性的影响来评估。
英文摘要
1510563(Stepp) & 1509791 (Koch Fager)This work will develop and evaluate a system to allow individuals with unintelligible speech due to severe paralysis to control a speech synthesizer that includes prosody (changes in the pitch, loudness, and duration in speech that convey meaning). This advancement to the synthetic speech and the ease of its control by users will facilitate improved functionality of clinical communication systems, thus improving the quality of life of users. Natural and intelligible speech production in these individuals will increase their ability to participate actively in society and empower them to self-advocate for their own medical management.The research objective of this proposal is to test the hypothesis that providing users of alternative and augmentative communication (AAC) with a method for prosodic control will result in speech synthesis that is more natural to listeners and provides greater function to users. Up to 1.2% of the population is unable to meet daily communication needs using typical speech due to stroke or other neurological injury, requiring AAC to meet their communication needs. Their quality of life is strongly dependent on access to this communication, both for social interaction as well as to relay information about urgent medical needs. The most advanced AAC devices incorporate speech synthesis, allowing the users to communicate orally with others. However, the resulting synthetic speech is both unnatural and difficult for others to understand, and is often described as "robotic". Specifically, synthetic speech does not vary in pitch, loudness, or rhythm, the prosodic features utilized in typical speech to relay emotional state, utterance form (statement vs. question), irony, and emphasis.Asking AAC users to control each of these dimensions individually would result in an intractably slow and complex system, an unacceptable burden for individuals who already have considerably reduced communication rates. Instead, this project will leverage the fact that typical speech predictably uses these prosodic markers (pitch, loudness, rhythm) in concert. A novel AAC interface will be developed to allow users to modify the overall "stress" of synthetic speech output as a single dimension, in order to provide easily controlled, natural, and intelligible speech synthesis. The co-PIs will use their combined expertise in speech technology, clinical application of AAC, and real-time control of human-machine-interfaces to enable essential advancements in AAC technology to achieve three goals. In Research Goal 1, a multi-stress speech bank for concatenative speech synthesis will be created via a novel interactive procedure in which speech productions of healthy speakers are "misunderstood", thus prompting speakers to naturally emphasize specific target sounds in their repeated responses. This will result in a bank of triphones (sounds with a specific left and right context, based on surrounding sounds) with all potential combinations of sounds and stresses. Research Goal 2 is to develop an AAC interface that allows users to select phonemes (individual sounds of speech) using two-dimensional cursor control (e.g., head-tracking, eye-tracking) in which the stress of individual phonemes will be based on cursor dwell time. In Research Goal 3, the functionality of the AAC interface will be evaluated by testing its effect on the naturalness of communicative interactions.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金