Body Structure Aware Deep Crowd Counting

Body Structure Aware Deep Crowd Counting
复制标题

身体结构感知深度人群计数

DOI:
10.1109/tip.2017.2740160
复制
发表时间:
2018-03-01
影响因子:
10.6
通讯作者:
Han, Junwei
Han, Junwei
中科院分区:
计算机科学1区
文献类型:
--
作者:
Huang, Siyu;Li, Xi;Han, Junwei

文献摘要

被引文献

相似文献

人群计数是一项具有挑战性的任务,主要是由于密集人群中的严重闭塞。本文旨在从语义建模的角度,以更广泛的视野来解决人群计数。从本质上讲,人群计数是一个行人语义分析的任务,涉及三个关键因素:行人,头部,他们的上下文结构。人体不同部位的信息是判断某个位置是否存在人的重要线索。现有方法通常从直接建模整个身体或头部的视觉特性的角度来执行人群计数,而没有显式地捕获对人群计数至关重要的复合身体部位语义结构信息。在我们的方法中,我们首先制定的语义场景模型的人群计数的关键因素。然后,我们将人群计数问题转化为多任务学习问题,将语义场景模型转化为不同的子任务。最后,使用深度卷积神经网络来学习统一方案中的子任务。我们的方法编码的语义性质的人群计数,并提供了一种新的解决方案,行人语义分析。在实验中,我们的方法优于国家的最先进的方法在四个基准人群计数数据集。语义结构信息被证明是一个有效的线索在场景中的人群计数。
Crowd counting is a challenging task, mainly due to the severe occlusions among dense crowds. This paper aims to take a broader view to address crowd counting from the perspective of semantic modeling. In essence, crowd counting is a task of pedestrian semantic analysis involving three key factors: pedestrians, heads, and their context structure. The information of different body parts is an important cue to help us judge whether there exists a person at a certain position. Existing methods usually perform crowd counting from the perspective of directly modeling the visual properties of either the whole body or the heads only, without explicitly capturing the composite body-part semantic structure information that is crucial for crowd counting. In our approach, we first formulate the key factors of crowd counting as semantic scene models. Then, we convert the crowd counting problem into a multi-task learning problem, such that the semantic scene models are turned into different sub-tasks. Finally, the deep convolutional neural networks are used to learn the sub-tasks in a unified scheme. Our approach encodes the semantic nature of crowd counting and provides a novel solution in terms of pedestrian semantic analysis. In experiments, our approach outperforms the state-of-the-art methods on four benchmark crowd counting data sets. The semantic structure information is demonstrated to be an effective cue in scene of crowd counting.