An Engineering View on Emotions and Speech: From Analysis and Predictive Models to Responsible Human-Centered Applications

An Engineering View on Emotions and Speech: From Analysis and Predictive Models to Responsible Human-Centered Applications
复制标题

DOI:
10.1109/jproc.2023.3276209
复制
发表时间:
2023-10
影响因子:
20.6
通讯作者:
Shrikanth S. Narayanan
Shrikanth S. Narayanan
中科院分区:
计算机科学1区
文献类型:
--
作者:
Shrikanth S. Narayanan

文献摘要

相似文献

物联网技术的大幅增长和智能手机设备的普及增加了公众和行业对语音情感识别(SER)技术的关注。然而,概念,技术和社会挑战限制了这些技术在各个领域的广泛采用,包括医疗保健和教育。由于人类情感固有的复杂性和主观性,难以在高时间分辨率下获得可靠的标签,以及混淆真实的生活中情感表达的各种上下文和环境因素,当自动情感识别系统被称为“在野外”发挥作用时,这些挑战被放大。此外,社会和道德挑战阻碍了这些技术的广泛接受和采用,公众对用户隐私,公平性和可解释性提出了质疑。本文简要回顾了情感语音处理的历史,概述了当前最先进的SER方法,并讨论了算法方法,使这些技术可供所有人使用,最大限度地发挥其优势,并导致负责任的以人为本的计算应用。
The substantial growth of Internet-of-Things technology and the ubiquity of smartphone devices has increased the public and industry focus on speech emotion recognition (SER) technologies. Yet, conceptual, technical, and societal challenges restrict the wide adoption of these technologies in various domains, including, healthcare, and education. These challenges are amplified when automated emotion recognition systems are called to function “in-the-wild” due to the inherent complexity and subjectivity of human emotion, the difficulty of obtaining reliable labels at high temporal resolution, and the diverse contextual and environmental factors that confound the expression of emotion in real life. In addition, societal and ethical challenges hamper the wide acceptance and adoption of these technologies, with the public raising questions about user privacy, fairness, and explainability. This article briefly reviews the history of affective speech processing, provides an overview of current state-of-the-art approaches to SER, and discusses algorithmic approaches to render these technologies accessible to all, maximizing their benefits and leading to responsible human-centered computing applications.