Unsupervised Learning of Image Segmentation Based on Differentiable Feature Clustering

Unsupervised Learning of Image Segmentation Based on Differentiable Feature Clustering
复制标题

DOI:
10.1109/tip.2020.3011269
复制
发表时间:
2020-01-01
影响因子:
10.6
通讯作者:
Tanaka, Masayuki
Tanaka, Masayuki
中科院分区:
计算机科学1区
文献类型:
--
作者:
Kim, Wonjik;Kanezaki, Asako;Tanaka, Masayuki

文献摘要

被引文献

相似文献

本文研究了卷积神经网络(cnn)在无监督图像分割中的应用。与监督图像分割类似,本文提出的CNN为像素分配标签,表示像素所属的聚类。然而,在无监督图像分割中,没有预先指定训练图像或像素的地面真值标签。因此,一旦输入目标图像,就联合优化像素标签和特征表示,并通过梯度下降更新其参数。在本文提出的方法中,标签预测和网络参数学习交替迭代,以满足以下条件:(a)相似特征的像素分配相同的标签,(b)空间连续的像素分配相同的标签,(c)唯一标签的数量要大。虽然这些标准是不相容的,但所提出的方法最大限度地减少了相似性损失和空间连续性损失的组合,以找到一个合理的标签分配解决方案,很好地平衡了上述标准。这项研究的贡献有四方面。首先,我们提出了一种新的端到端无监督图像分割网络,该网络由归一化和用于可微聚类的argmax函数组成。其次,我们引入了一个空间连续性损失函数,减轻了以往工作中固定段边界的局限性。第三,我们提出了一种以涂鸦作为用户输入的分割方法的扩展,该方法在保持效率的同时显示出比现有方法更好的准确率。最后,我们介绍了该方法的另一种扩展:使用使用少量参考图像预训练的网络进行未见图像分割,而无需重新训练网络。在多个图像分割基准数据集上验证了该方法的有效性。
The usage of convolutional neural networks (CNNs) for unsupervised image segmentation was investigated in this study. Similar to supervised image segmentation, the proposed CNN assigns labels to pixels that denote the cluster to which the pixel belongs. In unsupervised image segmentation, however, no training images or ground truth labels of pixels are specified beforehand. Therefore, once a target image is input, the pixel labels and feature representations are jointly optimized, and their parameters are updated by the gradient descent. In the proposed approach, label prediction and network parameter learning are alternately iterated to meet the following criteria: (a) pixels of similar features should be assigned the same label, (b) spatially continuous pixels should be assigned the same label, and (c) the number of unique labels should be large. Although these criteria are incompatible, the proposed approach minimizes the combination of similarity loss and spatial continuity loss to find a plausible solution of label assignment that balances the aforementioned criteria well. The contributions of this study are four-fold. First, we propose a novel end-to-end network of unsupervised image segmentation that consists of normalization and an argmax function for differentiable clustering. Second, we introduce a spatial continuity loss function that mitigates the limitations of fixed segment boundaries possessed by previous work. Third, we present an extension of the proposed method for segmentation with scribbles as user input, which showed better accuracy than existing methods while maintaining efficiency. Finally, we introduce another extension of the proposed method: unseen image segmentation by using networks pre-trained with a few reference images without re-training the networks. The effectiveness of the proposed approach was examined on several benchmark datasets of image segmentation.