The SEMAINE Database: Annotated Multimodal Records of Emotionally Colored Conversations between a Person and a Limited Agent

The SEMAINE Database: Annotated Multimodal Records of Emotionally Colored Conversations between a Person and a Limited Agent
复制标题

DOI:
10.1109/t-affc.2011.20
复制
发表时间:
2012-01-01
影响因子:
11.2
通讯作者:
Schroeder, Marc
Schroeder, Marc
中科院分区:
计算机科学2区
文献类型:
--
作者:
McKeown, Gary;Valstar, Michel;Schroeder, Marc

文献摘要

被引文献

相似文献

SEMAINE创建了一个大型视听数据库,作为构建敏感人工智能(SAL)代理的迭代方法的一部分,可以让一个人参与持续的,情绪化的对话。用于构建代理的数据来自用户和模拟SAL代理的“操作员”之间的交互,在不同的配置中:固体SAL(设计为使操作员显示适当的非语言行为)和半自动SAL(设计为使用户的体验近似于与机器交互)。然后,我们记录了用户与开发的系统,自动SAL的互动,比较最具沟通能力的版本,减少非语言技能的版本。高质量的录音由五个高分辨率、高帧率的摄像机和四个麦克风同步录制。录音共有150名参与者,共959次与SAL人物的对话,每次持续约5分钟。固体SAL记录转录和广泛的注释:6-8个评分员每个剪辑跟踪五个情感维度和27个相关类别。其他场景也被标记在相同的模式上,但不那么完整。其他信息包括对所选提取物的FACS注释,笑声,笑声和颤抖的识别,以及用户与自动系统的互动程度。这些材料可通过网络数据库查阅。
SEMAINE has created a large audiovisual database as a part of an iterative approach to building Sensitive Artificial Listener (SAL) agents that can engage a person in a sustained, emotionally colored conversation. Data used to build the agents came from interactions between users and an "operator" simulating a SAL agent, in different configurations: Solid SAL (designed so that operators displayed an appropriate nonverbal behavior) and Semi-automatic SAL (designed so that users' experience approximated interacting with a machine). We then recorded user interactions with the developed system, Automatic SAL, comparing the most communicatively competent version to versions with reduced nonverbal skills. High quality recording was provided by five high-resolution, high-framerate cameras, and four microphones, recorded synchronously. Recordings total 150 participants, for a total of 959 conversations with individual SAL characters, lasting approximately 5 minutes each. Solid SAL recordings are transcribed and extensively annotated: 6-8 raters per clip traced five affective dimensions and 27 associated categories. Other scenarios are labeled on the same pattern, but less fully. Additional information includes FACS annotation on selected extracts, identification of laughs, nods, and shakes, and measures of user engagement with the automatic system. The material is available through a web-accessible database.