Group Split and Merge Prediction With 3D Convolutional Networks

Group Split and Merge Prediction With 3D Convolutional Networks
复制标题

DOI:
10.1109/lra.2020.2969947
复制
发表时间:
2020-04
影响因子:
5.2
通讯作者:
Allan Wang;Aaron Steinfeld
Allan Wang;Aaron Steinfeld
中科院分区:
计算机科学2区
文献类型:
--
作者:
Allan Wang;Aaron Steinfeld

文献摘要

相似文献

人群中的移动的机器人往往由于对行人行为评估不足而具有有限的导航能力。我们通过预测多人组中的分裂和合并来加强这种能力。成功的预测应该会带来更有效的规划,同时也会增加人类对机器人行为的接受度。我们采取了一种新的方法,制定这作为一个视频预测问题,其中组分裂或合并的预测给定的历史的几何社会群体的形状变换。我们从视频相关任务的3D卷积模型的成功中汲取灵感。通过将时间维度视为空间维度,修改后的C3D模型成功地捕获了执行预测任务所需的时间特征。我们在几个数据集上展示了性能,并分析了向其他设置的传输能力。虽然目前用于跟踪人体运动的方法没有明确设计用于此任务,但我们的方法在预测分裂和合并的发生方面表现得更好。我们还从模型的学习特征中得出人类的解释。
Mobile robots in crowds often have limited navigation capability due to insufficient evaluation of pedestrian behavior. We strengthen this capability by predicting splits and merges in multi-person groups. Successful predictions should lead to more efficient planning while also increasing human acceptance of robot behavior. We take a novel approach by formulating this as a video prediction problem, where group splits or merges are predicted given a history of geometric social group shape transformations. We take inspiration from the success of 3D convolution models for video-related tasks. By treating the temporal dimension as a spatial dimension, a modified C3D model successfully captures the temporal features required to perform the prediction task. We demonstrate performance on several datasets and analyze transfer ability to other settings. While current approaches for tracking human motion are not explicitly designed for this task, our approach performs significantly better at predicting the occurrence of splits and merges. We also draw human interpretations from the model's learned features.