Learning Semisupervised Multilabel Fully Convolutional Network for Hierarchical Object Parsing

Learning Semisupervised Multilabel Fully Convolutional Network for Hierarchical Object Parsing
复制标题

DOI:
10.1109/tnnls.2019.2931183
复制
发表时间:
2020-07-01
影响因子:
10.4
通讯作者:
Yan, Shuicheng
Yan, Shuicheng
中科院分区:
计算机科学1区
文献类型:
--
作者:
Liu, Xiaobai;Xu, Qian;Yan, Shuicheng

文献摘要

相似文献

提出了一种用于图像分层对象分析的半监督多标签全卷积网络(FCN)。我们将每个对象部分(例如眼睛和头部)视为一个类标签,并学习将每个图像像素分配给多个连贯的部分标签。与之前将部分标签视为独立类的方法不同,我们的方法显式地模拟了物体部分之间的内部关系,例如,眼睛得分高的像素也应该在头部得分高。这种关系直接反映了语义空间的结构,因此在学习深度表示时应予以尊重。我们通过在标记和未标记图像上引入多标签softmax损失函数并使用两个成对排序约束对其进行正则化来实现这一目标。第一个约束是基于一个流形假设,即视觉上和空间上彼此接近的图像像素应该协同分类为相同的部分标签。另一个约束用于强制没有像素从多个语义上相互冲突的标签中获得显著分数。所提出的损失函数对网络参数是可微的,因此可以用标准的随机梯度方法进行优化。我们在两个公共图像数据集上评估了所提出的分层对象分析方法,并将其与其他分析方法进行了比较。广泛的比较表明,我们的方法可以达到最先进的性能,同时使用比替代方法少50%的标记训练样本。
This article presents a semisupervised multilabel fully convolutional network (FCN) for hierarchical object parsing of images. We consider each object part (e.g., eye and head) as a class label and learn to assign every image pixel to multiple coherent part labels. Different from previous methods that consider part labels as independent classes, our method explicitly models the internal relationships between object parts, e.g., that a pixel highly scored for eyes should be highly scored for heads as well. Such relationships directly reflect the structure of the semantic space and thus should be respected while learning the deep representation. We achieve this objective by introducing a multilabel softmax loss function over both labeled and unlabeled images and regularizing it with two pairwise ranking constraints. The first constraint is based on a manifold assumption that image pixels being visually and spatially close to each other should be collaboratively classified as the same part label. The other constraint is used to enforce that no pixel receives significant scores from more than one label that are semantically conflicting with each other. The proposed loss function is differentiable with respect to network parameters and hence can be optimized by standard stochastic gradient methods. We evaluate the proposed method on two public image data sets for hierarchical object parsing and compare it with the alternative parsing methods. Extensive comparisons showed that our method can achieve state-of-the-art performance while using 50% less labeled training samples than the alternatives.