RI: Medium: Collaborative Research: Understanding Individual-Level Speech Variability: From Novel Articulatory Data to Robust Speaker Recognition
RI: Medium: Collaborative Research: Understanding Individual-Level Speech Variability: From Novel Articulatory Data to Robust Speaker Recognition
批准号:
1514544
负责人:
Shrikanth Narayanan
金额:
$119.95万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-09-01 至 2020-08-31
中文摘要
说话是人类独有的能力。声道是人类普遍使用的乐器,在产生言语时表现出极大的灵巧和技巧,以传达丰富的语言和副语言信息。该项目将使人们能够从根本上了解由于身体发声乐器的形状和大小不同,个人的语音清晰度是如何不同的。了解人们在语音产生方面的不同之处有助于创造出改进的自动说话人识别技术,这项技术对国家安全很重要。该项目可以为所有人口成员,包括儿童、老年人和以某种语言为非母语的人提供强大的基于语音的访问技术的设计提供信息。该项目的成果还有助于更好地了解和治疗人类语言清晰度受到影响的疾病(例如唇腭裂)、疾病(例如头颈癌、呼吸暂停)或损伤。来自200个人的新的成像数据,以及由跨学科团队创建的相关工具、注释和解释将与科学界广泛共享。该项目将为学生提供一个独特的综合语音科技研究培训机会。这个项目的首要目标是促进对声道形态和语音清晰度如何相互作用的科学理解,并解释说话者语音信号特性的变化和不变方面。科学上特别感兴趣的是个体在存在结构差异的情况下为实现语音对等而采取的发音策略的性质。同样令人感兴趣的是声道形态差异的哪些方面以及如何反映在声学语音信号中,以及这些差异是否可以从语音声学中估计出来。这一目标的一个关键部分是创建正向和反向计算模型,将声道细节与语音声学联系起来,以揭示单个说话人的差异,并为稳健的说话人识别技术的设计提供信息。这个项目超越了最先进的方法,专注于使用新的成像技术和计算建模来直接调查动态的人类声道,以阐明声道结构中说话人之间的可变性,以及实现语言发音的策略。使用我们帮助开发的具有整个运动声道的高空间分辨率的新型磁共振成像(具有出色的时间分辨率的动态实时2D和加速的体积3D),该项目将收集并量化覆盖北美主要方言地区的160名美国本土英语和40名非母语人士的语音产生的时空细节。研究结构(形状和大小)和功能(声道成形的动力学及其声学后果)之间相互作用的实验、理论和方法方法可以带来新的理论进步,改进说话人普遍使用的语言单位的语音特征。它还提供了解释个人特定语音模式的能力,这些模式可以提高对科学基础的理解,并创建稳健的自动说话人识别技术,不仅能够通过采用新颖的说话人相关特征来确定两个说话者是不同的,而且通过分析生物启发的结构和发音细节,还能够确定他们如何以及为什么不同。
英文摘要
Speech is a unique human capability. The vocal tract is the universal human instrument played with great dexterity and skill in the production of speech to convey rich linguistic and paralinguistic information. The project will enable fundamental understanding of how individuals differ in their speech articulation due to differences in shape and size of their physical vocal instrument. Knowledge of how people differ in their speech production can help create improved automatic speaker recognition, technologies important for national security. The project can inform design of technologies for robust speech-based access for all members of the population, including children, the elderly, and non-native speakers of a language. Results from the project can also assist in better understanding and treating disorders (e.g., cleft lip/palate), illness (e.g., head and neck cancer, apnea) or injury where human speech articulation is affected. The novel imaging data from 200 individuals, and associated tools, annotations and interpretations created by the interdisciplinary team will be shared broadly with the scientific community. The project will provide a unique research training opportunity for students in integrated speech science and technology. The overarching goal of this project is to advance scientific understanding of how vocal tract morphology and speech articulation interact and explain the variant and invariant aspects of speech signal properties across talkers. Of particular scientific interest is the nature of articulatory strategies adopted by individuals in the presence of structural differences across them to achieve phonetic equivalence. Equally of interest are what aspects of, and how, vocal tract morphological differences are reflected in the acoustic speech signal, and if those differences can be estimated from speech acoustics. A crucial part of this goal is to create forward and inverse computational models that relate vocal tract details to speech acoustics toward shedding light on individual speaker differences and informing design of robust speaker recognition technologies. This project goes beyond state-of-the-art methods by focusing on direct investigation of the dynamic human vocal tract using novel imaging techniques and computational modeling to illuminate inter-speaker variability in vocal tract structure, as well as the strategies by which linguistic articulation is implemented. Using novel Magnetic Resonance Imaging with superior spatial resolution of the entire moving vocal tract that we helped develop (dynamic realtime 2D with excellent temporal resolution and accelerated volumetric 3D), the project will gather and quantify spatio-temporal details of speech production from 160 native American English covering the major dialectal regions of North America and 40 non-native speakers. The experimental, theoretical, and methodological approaches investigating the interplay between structure (shape and size) and function (dynamics of vocal-tract shaping and its acoustic consequences) can lead to new theoretical advances with improved phonetic characterizations of linguistic units that are general across speakers. It also offers the ability to explain individual specific speech patterns that can improve both understanding the scientific underpinning and creating robust automatic speaker recognition technology, enabling to determine not only that two talkers are different by the adoption of novel speaker dependent features, but also how and why they differ, by analyzing biologically-inspired details of structure and articulation.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
RI Core: Medium: Structured variability in vocal tract articulation dynamics in speech
-
批准号:2311676
-
项目类别:Standard Grant
-
资助金额:$120.0万
-
财政年份:2023
-
负责人:Shrikanth Narayanan
-
依托单位:
RI: Small: Speaker-Specific Articulatory Strategies
-
批准号:1908865
-
项目类别:Continuing Grant
-
资助金额:$47.4万
-
财政年份:2019
-
负责人:Shrikanth Narayanan
-
依托单位:
Be a Scientist!
-
批准号:1008372
-
项目类别:Continuing Grant
-
资助金额:$41.89万
-
财政年份:2010
-
负责人:Shrikanth Narayanan
-
依托单位:
Collaborative Research: Computational Behavioral Science: Modeling, Analysis, and Visualization of Social and Communicative Behavior
-
批准号:1029373
-
项目类别:Continuing Grant
-
资助金额:$150.0万
-
财政年份:2010
-
负责人:Shrikanth Narayanan
-
依托单位:
RI: Large: An Integrated Approach to Creating Context Enriched Speech Translation Systems
-
批准号:0911009
-
项目类别:Continuing Grant
-
资助金额:$220.0万
-
财政年份:2009
-
负责人:Shrikanth Narayanan
-
依托单位:
SGER: Exploring Emotional Vocal Productions Through the Use of Real-Time Magnetic Resonance Imaging
-
批准号:0844243
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:2008
-
负责人:Shrikanth Narayanan
-
依托单位:
Collaborative Research: Modeling Creative and Emotive Improvisation in Theater Performance
-
批准号:0757414
-
项目类别:Standard Grant
-
资助金额:$42.16万
-
财政年份:2008
-
负责人:Shrikanth Narayanan
-
依托单位:
CAREER: Modeling and Optimizing User-Centric Mixed-Initiative Spoken Dialog Systems
-
批准号:0238514
-
项目类别:Continuing Grant
-
资助金额:$50.0万
-
财政年份:2003
-
负责人:Shrikanth Narayanan
-
依托单位:
IERI: Collaborative Research: Automating Early Assessment of Academic Standards for Very Young Native and Non-Native Speakers of American English
-
批准号:0326228
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:2003
-
负责人:Shrikanth Narayanan
-
依托单位:
ITR: A User-centric Content-based Approach to Indexing, Query and Retrieval of Music through Signal Processing and Knowledge-based Methods
-
批准号:0219912
-
项目类别:Continuing Grant
-
资助金额:$50.0万
-
财政年份:2002
-
负责人:Shrikanth Narayanan
-
依托单位:
海外基金