Deep Representation Learning for Affective Speech Signal Analysis and Processing: Preventing unwanted signal disparities

Deep Representation Learning for Affective Speech Signal Analysis and Processing: Preventing unwanted signal disparities
复制标题

DOI:
10.1109/msp.2021.3105939
复制
发表时间:
2021-11
影响因子:
14.9
通讯作者:
Chi-Chun Lee;K. Sridhar;Jeng-Lin Li;Wei-Cheng Lin;Bo-Hao Su;C. Busso
Chi-Chun Lee;K. Sridhar;Jeng-Lin Li;Wei-Cheng Lin;Bo-Hao Su;C. Busso
中科院分区:
工程技术1区
文献类型:
--
作者:
Chi-Chun Lee;K. Sridhar;Jeng-Lin Li;Wei-Cheng Lin;Bo-Hao Su;C. Busso

文献摘要

被引文献

相似文献

语音情感识别(SER)是一个重要的研究领域,它直接影响着我们日常生活中的应用,涉及教育、医疗保健、安全和国防、娱乐和人机交互等多个领域。许多其他语音信号建模任务的进步,如自动语音识别、文本到语音合成和说话人识别,导致了当前基于语音的技术的扩散。将SER解决方案集成到现有和未来的系统中可以将这些基于语音的解决方案提升到一个新的水平。语音是一种高度非平稳的信号,具有动态演化的时空模式。它通常需要一个复杂的表示建模框架来开发能够处理现实生活中的复杂性的算法。
Speech emotion recognition (SER) is an important research area, with direct impacts in applications of our daily lives, spanning education, health care, security and defense, entertainment, and human–computer interaction. The advances in many other speech signal modeling tasks, such as automatic speech recognition, text-to-speech synthesis, and speaker identification, have led to the current proliferation of speech-based technology. Incorporating SER solutions into existing and future systems can take these voice-based solutions to the next level. Speech is a highly nonstationary signal, with dynamically evolving spatial-temporal patterns. It often requires a sophisticated representation modeling framework to develop algorithms capable of handling real-life complexities.