SRNet: Improving Generalization in 3D Human Pose Estimation with a Split-and-Recombine Approach

SRNet: Improving Generalization in 3D Human Pose Estimation with a Split-and-Recombine Approach
复制标题

DOI:
10.1007/978-3-030-58568-6_30
复制
发表时间:
2020-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Ailing Zeng;Xiao Sun;Fuyang Huang;Minhao Liu;Qiang Xu;Stephen Lin
Ailing Zeng;Xiao Sun;Fuyang Huang;Minhao Liu;Qiang Xu;Stephen Lin
中科院分区:
其他
文献类型:
--
作者:
Ailing Zeng;Xiao Sun;Fuyang Huang;Minhao Liu;Qiang Xu;Stephen Lin

文献摘要

被引文献

相似文献

在训练集中罕见或不可见的人类姿势对于网络预测来说是一个挑战。与视觉识别中的长尾分布问题类似,这种姿势的例子数量很少,限制了网络对其建模的能力。有趣的是,局部分布受长尾问题的影响较小,即,罕见姿势内的局部关节配置可能出现在训练集中的其他姿势内,使得它们不那么罕见。我们建议利用这一事实更好地推广到罕见的和看不见的姿势。具体来说,我们的方法将身体分割成局部区域,并在单独的网络分支中处理它们,利用关节的位置主要取决于其局部身体区域内的关节的属性。通过将来自身体其余部分的全局上下文作为低维向量重新组合到每个分支中来维持全局一致性。随着相关性较低的身体区域的维数降低,网络分支内的训练集分布更接近地反映了局部姿势而不是全局身体姿势的统计数据,而不会牺牲对联合推理重要的信息。所提出的分裂和重组的方法,calledSRNet,可以很容易地适应单图像和时间模型,它导致了可观的改善,在罕见的和看不见的姿态的预测。
Human poses that are rare or unseen in a training set are challenging for a network to predict. Similar to the long-tailed distribution problem in visual recognition, the small number of examples for such poses limits the ability of networks to model them. Interestingly,localpose distributions suffer less from the long-tail problem, i.e., local joint configurations within a rare pose may appear within other poses in the training set, making them less rare. We propose to take advantage of this fact for better generalization to rare and unseen poses. To be specific, our method splits the body into local regions and processes them in separate network branches, utilizing the property that a joint’s position depends mainly on the joints within its local body region. Global coherence is maintained by recombining the global context from the rest of the body into each branch as a low-dimensional vector. With the reduced dimensionality of less relevant body areas, the training set distribution within network branches more closely reflects the statistics oflocalposes instead of global body poses, without sacrificing information important for joint inference. The proposed split-and-recombine approach, calledSRNet, can be easily adapted to both single-image and temporal models, and it leads to appreciable improvements in the prediction of rare and unseen poses.