Finding Archetypal Spaces for Data Using Neural Networks

Finding Archetypal Spaces for Data Using Neural Networks
复制标题

使用神经网络寻找数据的原型空间

DOI:
--
复制
发表时间:
2019
期刊:
arXiv.org
影响因子:
--
通讯作者:
Smita Krishnaswamy
Smita Krishnaswamy
中科院分区:
--
文献类型:
--
作者:
D. V. Dijk;Daniel B. Burkhardt;Matthew Amodio;Alexander Tong;Guy Wolf;Smita Krishnaswamy

文献摘要

被引文献

相似文献

原型分析是一种因子分析,其中数据由凸多面体拟合,凸多面体的角是数据的“原型”,数据表示为这些原型点的凸组合。虽然原型分析已用于生物数据,但它尚未得到广泛采用,因为大多数数据都不能很好地适应环境空间中或标准数据转换后的凸多面体。我们提出了一种新的原型分析方法。我们不是直接在数据上或在特定的数据变换后拟合凸多面体,而是训练神经网络(AAnet)来学习数据最适合多面体的变换。我们在添加非线性的合成数据上验证了这种方法。在这里,AAnet 是唯一正确识别原型的方法。我们还在两个生物数据集上演示了 AAnet。在通过单细胞 RNA 测序测量的 T 细胞数据集中,AAnet 识别出与幼稚、记忆和细胞毒性 T 细胞相对应的几种原型状态。在肠道微生物组概况数据集中,AAnet 恢复了先前描述的微生物组状态并识别数据中的新极值。最后,我们证明 AAnet 具有生成特性,即使输入数据不均匀分布,我们也可以从数据几何中均匀采样。
Archetypal analysis is a type of factor analysis where data is fit by a convex polytope whose corners are "archetypes" of the data, with the data represented as a convex combination of these archetypal points. While archetypal analysis has been used on biological data, it has not achieved widespread adoption because most data are not well fit by a convex polytope in either the ambient space or after standard data transformations. We propose a new approach to archetypal analysis. Instead of fitting a convex polytope directly on data or after a specific data transformation, we train a neural network (AAnet) to learn a transformation under which the data can best fit into a polytope. We validate this approach on synthetic data where we add nonlinearity. Here, AAnet is the only method that correctly identifies the archetypes. We also demonstrate AAnet on two biological datasets. In a T cell dataset measured with single cell RNA-sequencing, AAnet identifies several archetypal states corresponding to naive, memory, and cytotoxic T cells. In a dataset of gut microbiome profiles, AAnet recovers both previously described microbiome states and identifies novel extrema in the data. Finally, we show that AAnet has generative properties allowing us to uniformly sample from the data geometry even when the input data is not uniformly distributed.