You-Do, I-Learn: Egocentric unsupervised discovery of objects and their modes of interaction towards video-based guidance

You-Do, I-Learn: Egocentric unsupervised discovery of objects and their modes of interaction towards video-based guidance
复制标题

DOI:
10.1016/j.cviu.2016.02.016
复制
发表时间:
2016-08-01
影响因子:
4.5
通讯作者:
Mayol-Cuevas, Walterio
Mayol-Cuevas, Walterio
中科院分区:
计算机科学3区
文献类型:
--
作者:
Damen, Dima;Leelasawassuk, Teesid;Mayol-Cuevas, Walterio

文献摘要

被引文献

相似文献

本文提出了一种无监督的方法,自动提取基于视频的指导对象的使用,从自我中心的视频和可穿戴的视线跟踪,从多个用户在执行任务时收集。该方法(i)发现任务相关对象,(ii)为每个对象建立模型,(iii)区分每个发现的对象被使用的不同方式,以及(iv)发现对象交互之间的依赖关系。这项工作使用外观,位置,运动和注意力进行调查,并使用每个相关功能的组合呈现结果。此外,一个在线的可扩展的方法,并比较离线的结果。本文提出了一种方法,用于选择一个合适的视频指南显示给新手用户,指示如何使用一个对象,纯粹由用户的目光触发。潜在辅助模式还可以基于所学习的对象交互序列来推荐接下来要使用的对象。该方法在各种日常任务中进行了测试,例如初始化打印机,准备咖啡和设置健身机。(C)2016 Elsevier Inc. All rights reserved.
This paper presents an unsupervised approach towards automatically extracting video-based guidance on object usage, from egocentric video and wearable gaze tracking, collected from multiple users while performing tasks. The approach (i) discovers task relevant objects, (ii) builds a model for each, (iii) distinguishes different ways in which each discovered object has been used and (iv) discovers the dependencies between object interactions. The work investigates using appearance, position, motion and attention, and presents results using each and a combination of relevant features. Moreover, an online scalable approach is presented and is compared to offline results. The paper proposes a method for selecting a suitable video guide to be displayed to a novice user indicating how to use an object, purely triggered by the user's gaze. The potential assistive mode can also recommend an object to be used next based on the learnt sequence of object interactions. The approach was tested on a variety of daily tasks such as initialising a printer, preparing a coffee and setting up a gym machine. (C) 2016 Elsevier Inc. All rights reserved.