Multi-Sensor Integration for Key-Frame Extraction From First-Person Videos

Multi-Sensor Integration for Key-Frame Extraction From First-Person Videos
复制标题

用于从第一人称视频中提取关键帧的多传感器集成

DOI:
10.1109/access.2020.3007150
复制
发表时间:
2020
期刊:
影响因子:
3.9
通讯作者:
Motoaki Kawanabe
Motoaki Kawanabe
中科院分区:
计算机科学3区
文献类型:
--
作者:
Yujie Li;Atsunori Kanemura;Hideki Asoh;Taiki Miyanishi;Motoaki Kawanabe

文献摘要

参考文献

相似文献

第一人称视觉(FPV)视频的关键帧提取是我们日常活动中选择重要场景、记忆深刻生活经历的核心技术。选择关键帧的困难是由于拍摄FPV视频时使用的头戴式摄像机造成的场景不稳定。由于头戴式摄像机往往会频繁抖动,因此FPV视频中的帧比第三人称视频(TPV)中的帧更嘈杂。然而,大多数现有的关键帧提取算法主要集中在处理TPV视频中的稳定场景。噪声FPV视频关键帧提取技术目前的技术发展还不成熟。此外,大多数关键帧提取算法主要使用来自FPV视频的视觉信息,尽管我们在日常活动中的视觉体验与人体运动有关。为了将FPV视频中动态变化场景的特征融入我们的方法中,将运动与视觉场景相结合是必不可少的。在本文中,我们提出了一种新的FPV视频关键帧提取方法,该方法利用多模态传感器信号通过典型相关分析(CCA)将多模态传感器信号投影到公共空间来降低噪声并检测显著活动。我们证明了提出的两种多传感器集成模型(基于稀疏的模型和基于图的模型)在公共空间上表现良好。在不同数据集上的实验结果表明,所提出的关键帧提取技术提高了提取的精度和对整个视频序列的覆盖。
Key-frame extraction for first-person vision (FPV) videos is a core technology for selecting important scenes and memorizing impressive life experiences in our daily activities. The difficulty of selecting key frames is the scene instability caused by head-mounted cameras used for capturing FPV videos. Because head-mounted cameras tend to frequently shake, the frames in an FPV video are noisier than those in a third-person vision (TPV) video. However, most existing algorithms for key-frame extraction mainly focus on handling the stable scenes in TPV videos. The technical development of key-frame extraction techniques for noisy FPV videos is currently immature. Moreover, most key-frame extraction algorithms mainly use visual information from FPV videos, even though our visual experience in daily activities is associated with human motions. To incorporate the features of dynamically changing scenes in FPV videos into our methods, integrating motions with visual scenes is essential. In this paper, we propose a novel key-frame extraction method for FPV videos that uses multi-modal sensor signals to reduce noise and detect salient activities via projecting multi-modal sensor signals onto a common space by canonical correlation analysis (CCA). We show that the two proposed multi-sensor integration models for key-frame extraction (a sparse-based model and a graph-based model) work well on the common space. The experimental results obtained using various datasets suggest that the proposed key-frame extraction techniques improve the precision of extraction and the coverage of entire video sequences.
DOI: 10.1007/s11263-013-0658-4
发表时间: 2014-01-01
影响因子: 19.5
作者:
Gong, Yunchao;Ke, Qifa;Lazebnik, Svetlana
通讯作者: Lazebnik, Svetlana
DOI: --
发表时间: 2005
期刊: --
影响因子: --
作者:
F. Bach;Michael I. Jordan
通讯作者: F. Bach;Michael I. Jordan
DOI: 10.1198/jasa.2008.s236
发表时间: 2008-06
影响因子: 3.7
作者:
T. Burr
通讯作者: T. Burr
DOI: 10.1109/tmm.2018.2818015
发表时间: 2018-03
影响因子: 7.3
作者:
Xiaohua Huang;Abhinav Dhall;R. Goecke;M. Pietikäinen;Guoying Zhao
通讯作者: Xiaohua Huang;Abhinav Dhall;R. Goecke;M. Pietikäinen;Guoying Zhao
DOI: 10.1109/tmm.2013.2291967
发表时间: 2014-02
影响因子: 7.3
作者:
Fan Chen;C. Vleeschouwer;A. Cavallaro
通讯作者: Fan Chen;C. Vleeschouwer;A. Cavallaro