Learning to Predict Diverse Human Motions from a Single Image via Mixture Density Networks

Learning to Predict Diverse Human Motions from a Single Image via Mixture Density Networks
复制标题

DOI:
10.1016/j.knosys.2022.109549
复制
发表时间:
2021-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Chunzhi Gu;Yan Zhao;Chao Zhang
Chunzhi Gu;Yan Zhao;Chao Zhang
中科院分区:
其他
文献类型:
--
作者:
Chunzhi Gu;Yan Zhao;Chao Zhang

文献摘要

相似文献

人体运动预测在计算机视觉中起着关键作用,通常需要过去的运动序列作为输入。然而,在实际应用中,实现完整且正确的过去运动序列的成本可能太高。在本文中,我们提出了一种利用混合密度网络(MDN)建模从更弱的条件(即单个图像)预测未来人体运动的新方法。与大多数现有的深度人体运动预测方法相反,MDN 的多模态性质使得能够生成多种未来运动假设,这很好地补偿了由单个输入和人体运动不确定性聚合的强随机模糊性。在设计损失函数时,我们进一步引入基于能量的公式,以灵活地对 MDN 的可学习参数施加先验损失,以保持运动相干性,并通过定制能量函数来提高预测精度。我们训练的模型直接将图像作为输入并生成满足给定条件的多个合理运动。对两个标准基准数据集的广泛实验证明了我们的方法在预测多样性和准确性方面的有效性。
Human motion prediction, which plays a key role in computer vision, generally requires a past motion sequence as input. However, in real applications, a complete and correct past motion sequence can be too expensive to achieve. In this paper, we propose a novel approach to predicting future human motions from a much weaker condition, i.e., a single image, with mixture density networks (MDN) modeling. Contrary to most existing deep human motion prediction approaches, the multimodal nature of MDN enables the generation of diverse future motion hypotheses, which well compensates for the strong stochastic ambiguity aggregated by the single input and human motion uncertainty. In designing the loss function, we further introduce the energy-based formulation to flexibly impose prior losses over the learnable parameters of MDN to maintain motion coherence as well as improve the prediction accuracy by customizing the energy functions. Our trained model directly takes an image as input and generates multiple plausible motions that satisfy the given condition. Extensive experiments on two standard benchmark datasets demonstrate the effectiveness of our method in terms of prediction diversity and accuracy.