BasisVAE: Translation-invariant feature-level clustering with Variational Autoencoders

BasisVAE: Translation-invariant feature-level clustering with Variational Autoencoders
复制标题

DOI:
--
复制
发表时间:
2020-03
期刊:
--
影响因子:
--
通讯作者:
Kaspar Märtens;C. Yau
Kaspar Märtens;C. Yau
中科院分区:
其他
文献类型:
--
作者:
Kaspar Märtens;C. Yau

文献摘要

相似文献

变分自动编码器(VAE)为非线性降维提供了一个灵活且可扩展的框架。然而,在应用领域,如基因组学的数据集通常是表格和高维的,黑盒方法降维不能提供足够的见解。常见的数据分析工作流还使用聚类技术来识别相似特征的组。这通常会导致一个两阶段的过程,但是,这将是可取的,以建立一个联合建模框架,同时降维和聚类的功能。在本文中,我们建议通过BasisVAE来实现这一目标:VAE和概率聚类先验的组合,它让我们学习作为解码器网络一部分的独热基函数表示。此外,对于不是所有特征都对齐的情况,我们开发了一个扩展来处理双线性不变基函数。我们展示了折叠变分推理方案如何为BasisVAE带来可扩展且高效的推理,并在各种玩具示例以及单细胞基因表达数据上进行了演示。
Variational Autoencoders (VAEs) provide a flexible and scalable framework for non-linear dimensionality reduction. However, in application domains such as genomics where data sets are typically tabular and high-dimensional, a black-box approach to dimensionality reduction does not provide sufficient insights. Common data analysis workflows additionally use clustering techniques to identify groups of similar features. This usually leads to a two-stage process, however, it would be desirable to construct a joint modelling framework for simultaneous dimensionality reduction and clustering of features. In this paper, we propose to achieve this through the BasisVAE: a combination of the VAE and a probabilistic clustering prior, which lets us learn a one-hot basis function representation as part of the decoder network. Furthermore, for scenarios where not all features are aligned, we develop an extension to handle translation-invariant basis functions. We show how a collapsed variational inference scheme leads to scalable and efficient inference for BasisVAE, demonstrated on various toy examples as well as on single-cell gene expression data.