Action recognition: From static datasets to moving robots

Action recognition: From static datasets to moving robots
复制标题

动作识别:从静态数据集到移动机器人

DOI:
10.1109/icra.2017.7989361
复制
发表时间:
2017
期刊:
2017 IEEE International Conference on Robotics and Automation (ICRA)
影响因子:
--
通讯作者:
Michael Milford
Michael Milford
中科院分区:
--
文献类型:
--
作者:
Fahimeh Rezazadegan;S. Shirazi;B. Upcroft;Michael Milford

文献摘要

被引文献

相似文献

深度学习模型在识别人类活动方面取得了最先进的性能,但往往依赖于利用典型计算机视觉数据集中的背景线索,这些数据集中主要有固定的摄像头。如果这些模型要用于现实世界环境中的自主机器人,它们必须被调整为独立于背景线索和相机运动效果而执行。为了应对这些挑战,我们提出了一种新的方法,该方法首先生成通用的动作区域建议,该建议具有在不受限制的视频中定位单个动作的良好潜力,而不考虑摄像机的运动,然后使用动作建议来提取和分类有效的形状和运动特征,并使用ConvNet框架。在一系列实验中,我们证明了通过在训练和测试期间积极提出动作区域,在基准测试中实现了最先进或更好的性能。我们在两个新的数据集上展示了我们的方法相对于最先进的数据集的性能;一个强调无关的背景,另一个强调相机的运动。我们还在一个异常行为检测场景中验证了我们的动作识别方法,以提高工作场所的安全性。结果验证了我们的方法有更高的成功率,这是因为我们的系统能够识别人类的行为,而不考虑环境和摄像机的运动。
Deep learning models have achieved state-of-the-art performance in recognizing human activities, but often rely on utilizing background cues present in typical computer vision datasets that predominantly have a stationary camera. If these models are to be employed by autonomous robots in real world environments, they must be adapted to perform independently of background cues and camera motion effects. To address these challenges, we propose a new method that firstly generates generic action region proposals with good potential to locate one human action in unconstrained videos regardless of camera motion and then uses action proposals to extract and classify effective shape and motion features by a ConvNet framework. In a range of experiments, we demonstrate that by actively proposing action regions during both training and testing, state-of-the-art or better performance is achieved on benchmarks. We show the outperformance of our approach compared to the state-of-the-art in two new datasets; one emphasizes on irrelevant background, the other highlights the camera motion. We also validate our action recognition method in an abnormal behavior detection scenario to improve workplace safety. The results verify a higher success rate for our method due to the ability of our system to recognize human actions regardless of environment and camera motion.