CIGMO: Categorical invariant representations in a deep generative framework

CIGMO: Categorical invariant representations in a deep generative framework
复制标题

CIGMO:深层生成框架中的分类不变表示

DOI:
10.48550/arxiv.2205.13758
复制
发表时间:
2022
期刊:
影响因子:
64.8
通讯作者:
H. Hosoya
H. Hosoya
中科院分区:
综合性期刊1区
文献类型:
--
作者:
H. Hosoya

文献摘要

参考文献

相似文献

一般对象图像的数据具有两种最常见的结构:(1)给定形状的每个对象可以在多个不同视图中呈现,以及(2)对象的形状可以以这样的方式进行分类:类别之间的形状多样性比类别内的形状多样性大得多。现有的深度生成模型通常可以捕获任一结构,但不能同时捕获两者。在这项工作中,我们引入了一种新颖的深度生成模型,称为 CIGMO,它可以学习表示图像数据中的类别、形状和视图因素。该模型由多个形状表示模块组成,每个模块专门针对特定类别并与视图表示分离,并且可以使用基于组的弱监督学习方法进行学习。通过实证研究,我们表明,尽管视图变化很大,我们的模型仍然可以有效地发现对象形状的类别,并定量地取代以前的各种方法,包括最先进的不变聚类算法。此外,我们表明,我们使用类别专业化的方法可以增强学习的形状表示,以更好地执行下游任务,例如一次性对象识别以及形状视图解开。
Data of general object images have two most common structures: (1) each object of a given shape can be rendered in multiple different views, and (2) shapes of objects can be categorized in such a way that the diversity of shapes is much larger across categories than within a category. Existing deep generative models can typically capture either structure, but not both. In this work, we introduce a novel deep generative model, called CIGMO, that can learn to represent category, shape, and view factors from image data. The model is comprised of multiple modules of shape representations that are each specialized to a particular category and disentangled from view representation, and can be learned using a group-based weakly supervised learning method. By empirical investigation, we show that our model can effectively discover categories of object shapes despite large view variation and quantitatively supersede various previous methods including the state-of-the-art invariant clustering algorithm. Further, we show that our approach using category-specialization can enhance the learned shape representation to better perform down-stream tasks such as one-shot object identification as well as shape-view disentanglement.
DOI: --
发表时间: --
期刊: --
影响因子: --
作者:
A R Simard;D. Soulet;G. Gowing;J. P. Julien;S. Rivest;B Ajami;J. L. Bennett;C. Krieger;W. Tetzlaff;F. M. Rossi;S. P. Sorokin;R. F. Hoyt;D. G. Blunt;N. McNelly;G. Hoeffel;X. H. Zong;R. Basu;H. Ketchum;W. Freiwald;Doris Y. Tsao
通讯作者: A R Simard;D. Soulet;G. Gowing;J. P. Julien;S. Rivest;B Ajami;J. L. Bennett;C. Krieger;W. Tetzlaff;F. M. Rossi;S. P. Sorokin;R. F. Hoyt;D. G. Blunt;N. McNelly;G. Hoeffel;X. H. Zong;R. Basu;H. Ketchum;W. Freiwald;Doris Y. Tsao