EAGER: Exploring the Use of Synthetic Speech as Reference Model to Detect Salient Emotional Segments in Speech
EAGER: Exploring the Use of Synthetic Speech as Reference Model to Detect Salient Emotional Segments in Speech
批准号:
1329659
负责人:
Carlos Busso
金额:
$5.93万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2013
资助国家:
美国
项目状态:
已结题
起止时间:
2013-03-15 至 2014-08-31
中文摘要
EArly探索性研究基金旨在从合成语音中创建中性参考模型,以对比语音信号的情感内容。情感理解是人类沟通的关键技能。出于这个原因,在设计和实现更符合用户需求的界面时,建模和识别情感是必不可少的。本研究从语言信息在时间上的非均匀性出发,旨在通过不同的声学特征来识别情感显著区域或焦点。该研究探索了一种基于合成语音的新方法,以建立表征中性语音中观察到的模式的参考模型。这些参考模型用于对比在语音信号的局部片段中观察到的情感信息。该研究建立了一个合成的语音信号,传达相同的词汇信息,并及时与数据库中的目标句子对齐。由于预计一个单一的合成语音将无法捕捉到中性语音中观察到的全部可变性,因此本研究探讨了产生不同中性合成实现的方法。在创建了一个与时间对齐的合成语音的平行语料库后,该研究探讨了合成语音如何很好地捕捉中性,非情感语音的声学模式和情感感知。然后,将来自数据库的目标信号与在合成信号族中观察到的特性进行比较。这项研究提出了一种新的方法来建立一个强大的情感识别系统,利用潜在的非均匀的表达行为的外部化过程。能够识别本地化情感片段的算法有可能改变情感计算领域目前使用的方法。而不是识别的情感内容的预分割的句子,问题被制定为一个检测范式,这是从应用的角度来看,有吸引力的。这些进步代表了行为分析和情感计算领域的变革性突破。所提出的模型和算法提供了许多见解,探索和扩展理论的语言学和人类行为学。 在为这项探索性研究建立了基础设施之后,将出现几种新的科学途径,这些途径将成为真正的创新进步,影响安全和国防,下一代高级用户界面,健康信息学和教育领域的应用。此外,科学方法丰富了本科生和研究生跨学科培训和指导的场所。
英文摘要
This EArly Grant for Exploratory Research aims to create neutral reference model from synthetic speech to contrast the emotional content of a speech signal. Emotional understanding is a crucial skill in human communication. For this reason, modeling and recognizing emotions is essential in the design and implementation of interfaces that are more in tune with the user's needs. Starting from the premise that paralinguistic information is non-uniformly conveyed across time, this study aims to identify emotionally prominent regions or focal points across various acoustic features. The study explores a novel approach based on synthetic speech to build reference models characterizing patterns observed in neutral speech. These reference models are used to contrast the emotional information observed in localized segments of a speech signal. The study builds a synthetic speech signal that conveys the same lexical information and is timely aligned with the target sentence in the database. Since it is expected that a single synthetic speech will not capture the full range of variability observed in neutral speech, the study explores approaches to produce different neutral synthetic realizations. After creating a parallel corpus with time-aligned synthetic speech, the study explores how well synthetic speech captures the acoustic patterns and emotional percepts of neutral, nonemotional speech. Then, a target signal from the database is compared with the properties observed across the family of synthesized signals. The study presents a novel approach to build a robust emotion recognition system that exploits the underlying nonuniform externalization process of expressive behaviors. Algorithms that able to identify localized emotional segments have the potential to shift the current approaches used in the area of affective computing. Instead of recognizing the emotional content of pre-segmented sentences, the problem is formulated as a detection paradigm, which is appealing from an application perspective. These advances represent a transformative breakthrough in the area of behavioral analysis and affective computing. The proposed models and algorithms provide numerous insights to explore and extend theories in linguistic and paralinguistic human behavior. Having established the base infrastructure for this exploratory research, several new scientific avenues will emerge that serve as truly innovative advancements that will impact applications in security and defense, next generation of advanced user interfaces, health informatics, and education. Furthermore, the scientific methods are enriching venues for interdisciplinary training and mentoring for undergraduate and graduate students.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
Lexical Dependent Emotion Detection Using Synthetic Speech Reference
使用合成语音参考进行词汇相关情绪检测
DOI:
10.1109/access.2019.2898353
发表时间:
2019
期刊:
IEEE Access
影响因子:
3.9
作者:
[Lotfian, Reza, Busso, Carlos]
通讯作者:
Busso, Carlos
CCRI: Medium: MSP-Podcast: Creating The Largest Speech Emotional Database By Leveraging Existing Naturalistic Recordings
-
批准号:2016719
-
项目类别:Standard Grant
-
资助金额:$107.54万
-
财政年份:2020
-
负责人:Carlos Busso
-
依托单位:
CRI: CI-P: Creating the Largest Speech Emotional Database by Leveraging Existing Naturalistic Recordings
-
批准号:1823166
-
项目类别:Standard Grant
-
资助金额:$9.94万
-
财政年份:2018
-
负责人:Carlos Busso
-
依托单位:
RI: Small: Integrative, Semantic-Aware, Speech-Driven Models for Believable Conversational Agents with Meaningful Behaviors
-
批准号:1718944
-
项目类别:Standard Grant
-
资助金额:$49.41万
-
财政年份:2017
-
负责人:Carlos Busso
-
依托单位:
FG 2015 Doctoral Consortium: Travel Support for Graduate Students
-
批准号:1540944
-
项目类别:Standard Grant
-
资助金额:$1.1万
-
财政年份:2015
-
负责人:Carlos Busso
-
依托单位:
CAREER: Advanced Knowledge Extraction of Affective Behaviors During Natural Human Interaction
-
批准号:1453781
-
项目类别:Continuing Grant
-
资助金额:$49.59万
-
财政年份:2015
-
负责人:Carlos Busso
-
依托单位:
WORKSHOP: Doctoral Consortium for the International Conference on Multimodal Interaction (ICMI 2013)
-
批准号:1346655
-
项目类别:Standard Grant
-
资助金额:$1.78万
-
财政年份:2013
-
负责人:Carlos Busso
-
依托单位:
RI: Small: Collaborative Research: Exploring Audiovisual Emotion Perception using Data-Driven Computational Modeling
-
批准号:1217104
-
项目类别:Continuing Grant
-
资助金额:$20.16万
-
财政年份:2012
-
负责人:Carlos Busso
-
依托单位:
Workshop: Doctoral Consortium at the 14th International Conference on Multimodal Interaction
-
批准号:1249319
-
项目类别:Standard Grant
-
资助金额:$1.46万
-
财政年份:2012
-
负责人:Carlos Busso
-
依托单位:
国内基金
海外基金
Exploring Changing Fertility Intentions in China
-
批准号:--
-
项目类别:外国学者研究基金
-
资助金额:--
-
批准年份:2024
-
负责人:MINHEE CHAE
-
依托单位:
Exploring the Intrinsic Mechanisms of CEO Turnover and Market
-
批准号:--
-
项目类别:外国学者研究基金
-
资助金额:--
-
批准年份:2024
-
负责人:HAOFEI Z
-
依托单位:
Exploring the Intrinsic Mechanisms of CEO Turnover and Market Reaction: An Explanation Based on Information Asymmetry
-
批准号:W2433169
-
项目类别:外国学者研究基金项目
-
资助金额:--
-
批准年份:2024
-
负责人:HAOFEI ZHANG
-
依托单位: