RI: Medium: Collaborative Research: Variance and Invariance in Voice Quality: Implications for Machine and Human Speaker Identification
RI: Medium: Collaborative Research: Variance and Invariance in Voice Quality: Implications for Machine and Human Speaker Identification
批准号:
1704167
负责人:
Abeer Alwan
金额:
$85.16万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-09-01 至 2023-08-31
中文摘要
点击翻译按钮获取中文摘要
英文摘要
A talker's voice quality conveys many kinds of information, including word and utterance prosody, emotional state, and personal identity. Variations in both the voice source and the vocal tract affect voice quality and there can be significant inter- and intra-talker variability. Understanding what aspects of a voice are talker-specific should aid in understanding the human limits in perceiving speaker differences and in developing better speaker identification (SID) algorithms. Despite technological advances, the performance of current SID systems remains far from perfect, and degrades significantly when the training and testing conditions are mismatched especially in terms of speech style (conversational versus read for example), speaker's emotional status, when the utterances are short, and when the task is text-independent. The key questions that the project aims to answer are: under normal daily life variability, how often does a talker sound less like him- or herself and more like someone else? Which acoustic properties account for speaker similarity? Can automatic speaker identification (SID) algorithms be improved by knowledge of which properties are important for human perception of speaker similarity?The project is a transformative one and helps better understand and model variance and invariance in voice quality. It will inform several important issues in human speech perception, especially in the area of talker similarity. Understanding what aspects of the source signal, if any, are talker-specific, should aid in developing better speaker identification and verification algorithms that are able to handle short utterances and are robust to varying affect and styles of speaking. A model of voice quality variations could also improve the naturalness of text-to-speech (TTS) systems. If it were known how much a person could change his or her voice quality without compromising their vocal identity, this knowledge could also inform medical rehab applications and forensics. A better understanding of voice quality will thus be of significant impact scientifically, and for engineering, forensic, and medical applications. The project has strong outreach and dissemination programs and fosters interdisciplinary activities in Electrical Engineering, Linguistics, and Speech and Hearing Science at UCLA and the Center of Excellence at JHU. It trains undergraduate and graduate students in important cross-disciplinary activities of technological and scientific significance. The results will be published in high-quality journals and presented at relevant international conferences. The research results - a set of databases, software tools, and publications will be disseminated freely.The project analyzes and discovers how the speech signal varies within and across talkers under circumstances that introduce variability in everyday life situations. Specifically, it investigates whether an individual talker's speech varies significantly across recording sessions and speech tasks. Most importantly, it examines how intra-talker variability from all these sources of variability compares with inter-talker variability. Understanding these issues requires a high-quality speech database with multiple voice samples from many talkers (in this case 200) which are collected, annotated, and distributed to other researchers. Acoustic analyses reveals inter- and intra-talker variability in the speech signal across different situations by generating a multi- dimensional acoustic profile of each talker that specifies the range of parameter values that are typical in the corpus for that talker, and the likelihood of deviations from that usual profile. Perceptual studies determine the extent to which parameter profiles predict perceived similarity, and how much variability in each parameter can be tolerated before talkers cease to sound like themselves. Insights from the acoustic and perceptual studies guide the development of robust text-dependent and text-independent SID algorithms that are anticipated to be robust to variations in affect, style, and for short utterances.
期刊论文(19)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.21437/interspeech.2020-2957
发表时间:
2020-08
期刊:
影响因子:
--
作者:
[Vijay Ravi;Ruchao Fan;Amber Afshan;Huanhua Lu;A. Alwan]
通讯作者:
Vijay Ravi;Ruchao Fan;Amber Afshan;Huanhua Lu;A. Alwan
DOI:
10.21437/interspeech.2020-3006
发表时间:
2020-08
期刊:
ArXiv
影响因子:
--
作者:
[Amber Afshan;Jinxi Guo;S. Park;Vijay Ravi;A. McCree;A. Alwan]
通讯作者:
Amber Afshan;Jinxi Guo;S. Park;Vijay Ravi;A. McCree;A. Alwan
Target and Non-target Speaker Discrimination by Humans and Machines
人类和机器对目标和非目标说话者的辨别
DOI:
10.1109/icassp.2019.8683362
发表时间:
2019
期刊:
IEEE ICASSP 2019
影响因子:
--
作者:
[Park, Soo Jin, Afshan, Amber, Kreiman, Jody, Yeung, Gary, Alwan, Abeer]
通讯作者:
Alwan, Abeer
Acoustic voice variation in spontaneous speech
自发言语中的声学语音变化
DOI:
10.1121/10.0011471
发表时间:
2022
期刊:
The Journal of the Acoustical Society of America
影响因子:
--
作者:
[Lee, Yoonjeong, Kreiman, Jody]
通讯作者:
Kreiman, Jody
Acoustic voice variation within and between speakers
说话者内部和说话者之间的声音变化
DOI:
10.1121/1.5125134
发表时间:
2019
期刊:
The Journal of the Acoustical Society of America
影响因子:
--
作者:
[Lee, Yoonjeong, Keating, Patricia, Kreiman, Jody]
通讯作者:
Kreiman, Jody
共 17 条
Collaborative Research: Improving speech technology for better learning outcomes: the case of AAE child speakers
-
批准号:2202585
-
项目类别:Standard Grant
-
资助金额:$31.89万
-
财政年份:2022
-
负责人:Abeer Alwan
-
依托单位:
Collaborative Research: RI: Small: From Ultrasound and MRI to articulatory and acoustic models of child speech development
-
批准号:2006979
-
项目类别:Standard Grant
-
资助金额:$23.0万
-
财政年份:2020
-
负责人:Abeer Alwan
-
依托单位:
Workshop for Undergraduate and MS Female Students in Speech Science and Technology
-
批准号:1745166
-
项目类别:Standard Grant
-
资助金额:$2.5万
-
财政年份:2017
-
负责人:Abeer Alwan
-
依托单位:
NRI: INT: COLLAB: Development, Deployment and Evaluation of Personalized Learning Companion Robots for Early Literacy and Language Learning
-
批准号:1734380
-
项目类别:Standard Grant
-
资助金额:$61.56万
-
财政年份:2017
-
负责人:Abeer Alwan
-
依托单位:
A Workshop for Junior Female Researchers in Speech Science and Technology
-
批准号:1637240
-
项目类别:Standard Grant
-
资助金额:$3.0万
-
财政年份:2016
-
负责人:Abeer Alwan
-
依托单位:
The Role of Speech Science in Developing Robust Speech Technology Applications
-
批准号:1543522
-
项目类别:Standard Grant
-
资助金额:$3.5万
-
财政年份:2015
-
负责人:Abeer Alwan
-
依托单位:
EAGER: Collaborative Research: Models of Child Speech
-
批准号:1551113
-
项目类别:Standard Grant
-
资助金额:$14.0万
-
财政年份:2015
-
负责人:Abeer Alwan
-
依托单位:
EAGER: Variance and Invariance in Voice Quality
-
批准号:1450992
-
项目类别:Standard Grant
-
资助金额:$20.0万
-
财政年份:2014
-
负责人:Abeer Alwan
-
依托单位:
EAGER: Collaborative Research: Towards Modeling Human Speech Confusions in Noise
-
批准号:1247809
-
项目类别:Standard Grant
-
资助金额:$10.0万
-
财政年份:2012
-
负责人:Abeer Alwan
-
依托单位:
RI: Small: A New Voice Source Model: From Glottal Areas to Better Speech Synthesis
-
批准号:1018863
-
项目类别:Continuing Grant
-
资助金额:$45.0万
-
财政年份:2010
-
负责人:Abeer Alwan
-
依托单位:
RI: Medium: Collaborative Research: The Effect of Subglottal Resonances on Machine and Human Speaker Normalization
-
批准号:0905381
-
项目类别:Standard Grant
-
资助金额:$63.97万
-
财政年份:2009
-
负责人:Abeer Alwan
-
依托单位:
Collaborative Research: IDBR: VoxNet--A deployable bioacoustic sensor network
-
批准号:0936454
-
项目类别:Continuing Grant
-
资助金额:$4.25万
-
财政年份:2008
-
负责人:Abeer Alwan
-
依托单位:
Collaborative Research: IDBR: VoxNet--A deployable bioacoustic sensor network
-
批准号:0754120
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2008
-
负责人:Abeer Alwan
-
依托单位:
Collaborative Research: Landmark-based Robust Speech Recognition using Prosody-guided Models of Speech Variability
-
批准号:0703805
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2007
-
负责人:Abeer Alwan
-
依托单位:
ITR-Collaborative Research: Development and Evaluation of a Hybrid Concatenative/Rule-Based Visual Speech Synthesis System
-
批准号:0312810
-
项目类别:Standard Grant
-
资助金额:$18.32万
-
财政年份:2003
-
负责人:Abeer Alwan
-
依托单位:
IERI Collaborative Research: Automating Early Assesment of Academic Standards for Very Young Native and Non-Native Speakers of American English
-
批准号:0326214
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:2003
-
负责人:Abeer Alwan
-
依托单位:
CAREER: From Imaging and Acoustic Data to Articulatory Synthesis
-
批准号:9503089
-
项目类别:Continuing Grant
-
资助金额:$13.94万
-
财政年份:1995
-
负责人:Abeer Alwan
-
依托单位:
RIA: A Model of Speech Perception in Noise
-
批准号:9309418
-
项目类别:Standard Grant
-
资助金额:$10.0万
-
财政年份:1993
-
负责人:Abeer Alwan
-
依托单位:
海外基金