A Multi-party Multi-modal Dataset for Focus of Visual Attention in Human-human and Human-robot Interaction

A Multi-party Multi-modal Dataset for Focus of Visual Attention in Human-human and Human-robot Interaction
复制标题

用于人与人以及人与机器人交互中视觉注意力焦点的多方多模态数据集

DOI:
--
复制
发表时间:
2016
期刊:
International Conference on Language Resources and Evaluation
影响因子:
--
通讯作者:
J. Beskow
J. Beskow
中科院分区:
--
文献类型:
--
作者:
Kalin Stefanov;J. Beskow

文献摘要

被引文献

相似文献

本文描述了一个数据收集设置和一个新记录的数据集。该数据集的主要目的是探索三种不同条件下人类视觉注意力焦点的模式-两个人参与与机器人的基于任务的交互;同样的两个人参与基于任务的交互,其中机器人被第三个人取代,以及自由的三方人类交互。该数据集包含两个部分- 6个持续时间约为3小时的会话和9个持续时间约为4.5小时的会话。数据集的两个部分都包含丰富的模态和记录的数据流-它们包括三个Kinect v2设备的流(颜色,深度,红外线,身体和面部数据),三个高质量音频流,三个高分辨率GoPro视频流,用于基于任务的交互的触摸数据和机器人的系统状态。此外,数据集的第二部分介绍了来自三台Tobii Pro Glasses 2眼动仪的数据流。所有交互的语言都是英语,所有数据流都在空间和时间上对齐。
This papers describes a data collection setup and a newly recorded dataset. The main purpose of this dataset is to explore patterns in the focus of visual attention of humans under three different conditions - two humans involved in task-based interaction with a robot; same two humans involved in task-based interaction where the robot is replaced by a third human, and a free three-party human interaction. The dataset contains two parts - 6 sessions with duration of approximately 3 hours and 9 sessions with duration of approximately 4.5 hours. Both parts of the dataset are rich in modalities and recorded data streams - they include the streams of three Kinect v2 devices (color, depth, infrared, body and face data), three high quality audio streams, three high resolution GoPro video streams, touch data for the task-based interactions and the system state of the robot. In addition, the second part of the dataset introduces the data streams from three Tobii Pro Glasses 2 eye trackers. The language of all interactions is English and all data streams are spatially and temporally aligned.