Deep Representation Learning for Affective Speech Signal Analysis and Processing: Preventing unwanted signal disparities
Deep Representation Learning for Affective Speech Signal Analysis and Processing: Preventing unwanted signal disparities
复制标题
DOI:
10.1109/msp.2021.3105939
复制
发表时间:
2021-11
影响因子:
14.9
通讯作者:
Chi-Chun Lee;K. Sridhar;Jeng-Lin Li;Wei-Cheng Lin;Bo-Hao Su;C. Busso
中科院分区:
文献类型:
--
作者:
Chi-Chun Lee;K. Sridhar;Jeng-Lin Li;Wei-Cheng Lin;Bo-Hao Su;C. Busso
Speech emotion recognition (SER) is an important research area, with direct impacts in applications of our daily lives, spanning education, health care, security and defense, entertainment, and human–computer interaction. The advances in many other speech signal modeling tasks, such as automatic speech recognition, text-to-speech synthesis, and speaker identification, have led to the current proliferation of speech-based technology. Incorporating SER solutions into existing and future systems can take these voice-based solutions to the next level. Speech is a highly nonstationary signal, with dynamically evolving spatial-temporal patterns. It often requires a sophisticated representation modeling framework to develop algorithms capable of handling real-life complexities.