A computational model for the automatic recognition of affect in speech

A computational model for the automatic recognition of affect in speech
复制标题

DOI:
--
复制
发表时间:
2004
期刊:
--
影响因子:
--
通讯作者:
Raul Fernandez;Rosalind W. Picard
Raul Fernandez;Rosalind W. Picard
中科院分区:
其他
文献类型:
--
作者:
Raul Fernandez;Rosalind W. Picard

文献摘要

被引文献

相似文献

口语除了作为语言结构和意义外化的主要载体外,还作为各种信息来源的载体,包括背景、年龄、性别、社会结构成员以及生理、病理和情感状态。这些信息来源不仅仅是辅助语言交流的主要目的:人类对语音信号中编码的各种非语言因素做出反应,塑造和调整他们的互动,以满足人际和社会协议。计算机科学、人工智能和计算语言学已经投入了大量积极的研究,旨在模拟语言词汇语义结构的产生和恢复。然而,很少有人关注的系统,模型和理解的语言和语言外的信息信号。随着人机交互的广度和性质升级到以前为人与人通信保留的水平,越来越需要赋予计算系统类似人类的能力,以促进交互并使其更自然。其中最重要的是人类能够对我们交流的情感内容进行推断。本文提出了一个从口语中提取韵律声学参数的情感限定词识别框架。有人认为,可以通过整合来自各种韵律时间尺度的声学参数、总结来自更局部化的信息(例如,音节级)到更全局的韵律现象(例如,话语水平)。在这个框架中,语音在结构上被建模为一个动态演变的层次模型,其中层次结构的级别由韵律选区确定,并包含根据动态系统演变的参数。声学参数已被选择来反映语音思想的四个主要组成部分,以反映语言学和情感特定的信息:语调,响度,节奏和语音质量。本文分别讨论了这些组件的贡献,并评估了完整的模型进行测试的数据集上的行为和自发语音感知注释的情感标签,并通过比较它对人类的性能基准。(副本可从麻省理工学院图书馆,RM。14-0551,剑桥,MA 02139-4307。电话:617-253-5668;传真:617-253-1690。)
Spoken language, in addition to serving as a primary vehicle for externalizing linguistic structures and meaning, acts as a carrier of various sources of information, including background, age, gender, membership in social structures, as well as physiological, pathological and emotional states. These sources of information are more than just ancillary to the main purpose of linguistic communication: Humans react to the various non-linguistic factors encoded in the speech signal, shaping and adjusting their interactions to satisfy interpersonal and social protocols. Computer science, artificial intelligence and computational linguistics have devoted much active research to systems that aim to model the production and recovery of linguistic lexico-semantic structures from speech. However, less attention has been devoted to systems that model and understand the paralinguistic and extralinguistic information in the signal. As the breadth and nature of human-computer interaction escalates to levels previously reserved for human-to-human communication, there is a growing need to endow computational systems with human-like abilities which facilitate the interaction and make it more natural. Of paramount importance amongst these is the human ability to make inferences regarding the affective content of our exchanges. This thesis proposes a framework for the recognition of affective qualifiers from prosodic-acoustic parameters extracted from spoken language. It is argued that modeling the affective prosodic variation of speech can be approached by integrating acoustic parameters from various prosodic time scales, summarizing information from more localized (e.g., syllable level) to more global prosodic phenomena (e.g., utterance level). In this framework speech is structurally modeled as a dynamically evolving hierarchical model in which levels of the hierarchy are determined by prosodic constituency and contain parameters that evolve according to dynamical systems. The acoustic parameters have been chosen to reflect four main components of speech thought to reflect paralinguistic and affect-specific information: intonation, loudness, rhythm and voice quality. The thesis addresses the contribution of each of these components separately, and evaluates the full model by testing it on datasets of acted and of spontaneous speech perceptually annotated with affective labels, and by comparing it against human performance benchmarks. (Copies available exclusively from MIT Libraries, Rm. 14-0551, Cambridge, MA 02139-4307. Ph. 617-253-5668; Fax 617-253-1690.)