The KIT Motion-Language Dataset

The KIT Motion-Language Dataset
复制标题

DOI:
10.1089/big.2016.0028
复制
发表时间:
2016-12-01
期刊:
影响因子:
4.6
通讯作者:
Asfour, Tamim
Asfour, Tamim
中科院分区:
计算机科学4区
文献类型:
--
作者:
Plappert, Matthias;Mandery, Christian;Asfour, Tamim

文献摘要

被引文献

相似文献

将人类运动和自然语言联系起来对于人类活动的语义表示的生成以及基于自然语言输入的机器人活动的生成非常有意义。然而,尽管该领域已经进行了多年的研究,但仍不存在标准化且公开的数据集来支持此类系统的开发和评估。因此,我们提出了卡尔斯鲁厄理工学院 (KIT) 运动语言数据集,该数据集庞大、开放且可扩展。我们聚合来自多个动作捕捉数据库的数据,并使用独立于捕捉系统或标记集的统一表示将它们包含在我们的数据集中,从而可以轻松处理数据,无论其来源如何。为了获得自然语言的运动注释,我们应用了众包方法和专门为此目的构建的基于网络的工具,即运动注释工具。我们彻底记录了注释过程本身,并讨论了我们用来保持注释者积极性的游戏化方法。我们进一步提出了一种新颖的方法,即基于困惑度的选择,它系统地选择运动以进行进一步注释,这些运动要么在我们的数据集中代表性不足,要么具有错误的注释。我们证明我们的方法可以缓解上述两个问题并确保系统的注释过程。我们对所得数据集的结构和内容进行了深入分析,截至 2016 年 10 月 10 日,该数据集包含 3911 个动作,总持续时间为 11.23 小时,以及 6278 个自然语言注释,包含 52,903 个单词。我们相信,这使我们的数据集成为一个绝佳的选择,可以在这一重要领域实现更加透明和可比的研究。
Linking human motion and natural language is of great interest for the generation of semantic representations of human activities as well as for the generation of robot activities based on natural language input. However, although there have been years of research in this area, no standardized and openly available data set exists to support the development and evaluation of such systems. We, therefore, propose the Karlsruhe Institute of Technology (KIT) Motion-Language Dataset, which is large, open, and extensible. We aggregate data from multiple motion capture databases and include them in our data set using a unified representation that is independent of the capture system or marker set, making it easy to work with the data regardless of its origin. To obtain motion annotations in natural language, we apply a crowd-sourcing approach and a web-based tool that was specifically build for this purpose, the Motion Annotation Tool. We thoroughly document the annotation process itself and discuss gamification methods that we used to keep annotators motivated. We further propose a novel method, perplexity-based selection, which systematically selects motions for further annotation that are either under-represented in our data set or that have erroneous annotations. We show that our method mitigates the two aforementioned problems and ensures a systematic annotation process. We provide an in-depth analysis of the structure and contents of our resulting data set, which, as of October 10, 2016, contains 3911 motions with a total duration of 11.23 hours and 6278 annotations in natural language that contain 52,903 words. We believe this makes our data set an excellent choice that enables more transparent and comparable research in this important area.