课题基金 / 基金详情

Collaborative Research: Estimating Articulatory Constriction Place and Timing from Speech Acoustics

Collaborative Research: Estimating Articulatory Constriction Place and Timing from Speech Acoustics
合作研究:从语音声学估计发音收缩位置和时间
批准号:
2141413
负责人:
Carol Espy-Wilson
金额:
$24.54万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-06-15 至 2025-11-30

项目摘要

项目成果

Carol Espy-Wilson的其他基金

相似基金

相关文献

中文摘要
翻译
这个合作项目的重点是一种使用语音录音来研究说话人发音习惯的新方法--也就是说话人如何系统地协调他们的嘴唇、下巴、舌头、声门和软腭的发音运动来产生单词和句子。这些发音习惯在不同的人之间,以及同一语言的不同语言和方言之间都不同,这是外国口音、言语障碍和说话风格的许多方面的原因。虽然以前对这些习惯的研究需要专门的设备来立即观察发音装置的运动,但这个项目的目的是开发和改进一种“语音反转”工具--即一种可以使用机器学习方法直接从声学语音信号中准确地恢复发音运动的工具。到目前为止,项目组开发的工具已经成功地恢复了舌头和嘴唇的运动;目前的项目将该工具的功能扩展到包括鼻音(软腭)和发声(声门)。将使用新收集的声学和发音语料库对扩展系统进行培训和验证,语料库来自说美国英语的人。这个语料库包括共同收集的音频、鼻音、语音和发音动作,将作为训练和评估完全训练的语音倒置系统的能力的“基本事实”。作为进一步的测试,我们将根据发音习惯模式与英语不同的语言使用者的基本真实数据对其进行测试。该项目的目标是开发和改进一种语音倒置工具,该工具可以读取语音的声学记录,并恢复发音运动的大小和时间的细节。该项目旨在通过训练专门的神经网络模型来实现这一目标,以将声信号的特征与单独获取的地面真实鼻腔与口腔流出信号和并发电声门图联系起来。训练数据来自以英语为母语的人;泛化的验证和测试包括加拿大人、法语人和俄语人的作品。当成功验证时,所产生的语音反转工具将有助于识别影响语音运动组织的医学问题,例如众所周知的构音障碍说话者的口腔/喉部计时中断。此外,结合发音能力的估计也有助于追踪抑郁症和精神分裂症等疾病引起的变化。更广泛地说,仅是快速、轻松地分析从录音中获得的发音动作的能力就有可能极大地改进自动语音识别(ASR)系统,并帮助学者、法医科学家和临床专业人员在农村或资源匮乏地区的野外条件下研究社区的语音,并帮助记录濒危语言。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This collaborative project focuses on a new approach for using speech recordings to study speaker pronunciation habits--that is, the way speakers systematically coordinate the articulatory movements of their lips, jaw, tongue, glottis and soft palate to produce words and sentences. These articulatory habits differ between individuals, and across languages and dialects of the same language, accounting for many aspects of foreign accent, speech disorders and speaking style. Whereas previous studies of these habits have required specialized equipment for the immediate observation of articulator movements, the aim of this project is to develop and improve a tool for "speech inversion"--that is, a tool that can accurately recover articulatory movements directly from the acoustic speech signal using machine learning methods. To date, the tool developed by the project team has successfully recovered movements of the tongue and lips; the current project extends the tool’s functionality to encompass nasality (soft palate) and voicing (glottis). Training and validation of the extended system will proceed using a newly collected corpus of acoustic and articulatory data drawn from speakers of American English. This corpus, comprising co-collected audio, nasal, voicing, and articulatory movement, will serve as 'ground truth' for training and assessing the capabilities of the fully trained speech inversion system. As a further test, we will test it against ground truth data from speakers of languages with patterns of articulatory habits known to differ from English.The goal of this project is to develop and refine a Speech Inversion Tool that 'reads' acoustic recordings of speech and 'recovers' details of the magnitude and timing of articulatory movements. The project aims to accomplish this goal by training specialized Neural Network models to relate features of the acoustic signal to separately acquired ground-truth nasal vs. oral outflow signals and concurrent electroglottography. Training data derives from native speakers of English; validation and tests for generalization include productions of speakers of Canadian French and Russian. When successfully validated, the resulting speech inversion tool will be useful for identifying medical issues that affect speech movement organization, such as the well-known disruption of oral/laryngeal timing in speakers with dysarthria. In addition, incorporating estimates of articulation may also aid in the tracking of changes resulting from medical conditions such as depression and schizophrenia. More generally, the ability to rapidly and easily analyze articulatory movements obtained from audio recordings alone has the potential substantially improve Automated Speech Recognition (ASR) systems, and to assist scholars, forensic scientists, and clinical professionals studying the speech of communities under field conditions in rural or under-resourced areas, and to help in the documentation of endangered languages.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
SCH: INT: Collaborative Research: Using Multi-Stage Learning to Prioritize Mental Health
  • 批准号:
    2124270
  • 项目类别:
    Standard Grant
  • 资助金额:
    $84.24万
  • 财政年份:
    2021
  • 负责人:
    Carol Espy-Wilson
  • 依托单位:
Speech for Robotics
  • 批准号:
    1941541
  • 项目类别:
    Standard Grant
  • 资助金额:
    $5.0万
  • 财政年份:
    2019
  • 负责人:
    Carol Espy-Wilson
  • 依托单位:
Collaborative Research: Effects of production variability on the acoustic consequences of coordinated articulatory gestures
  • 批准号:
    1436600
  • 项目类别:
    Standard Grant
  • 资助金额:
    $13.24万
  • 财政年份:
    2014
  • 负责人:
    Carol Espy-Wilson
  • 依托单位:
RI: Medium: Collaborative Research: Multilingual Gestural Models for Robust Language-Independent Speech Recognition
  • 批准号:
    1162525
  • 项目类别:
    Standard Grant
  • 资助金额:
    $23.49万
  • 财政年份:
    2012
  • 负责人:
    Carol Espy-Wilson
  • 依托单位:
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)