Autoencoder-based cluster ensembles for single-cell RNA-seq data analysis

Autoencoder-based cluster ensembles for single-cell RNA-seq data analysis
复制标题

DOI:
10.1186/s12859-019-3179-5
复制
发表时间:
2019-12-24
期刊:
影响因子:
3
通讯作者:
Yang, Pengyi
Yang, Pengyi
中科院分区:
生物学4区
文献类型:
--
作者:
Geddes, Thomas A.;Kim, Taiyun;Yang, Pengyi

文献摘要

被引文献

相似文献

背景:单细胞RNA测序(scRNA-seq)是一种变革性的技术,可以高精度地分析单个细胞的整体转录本。ScRNA-seq数据分析中的一项基本任务是从复杂的样本或实验中描绘的组织中识别细胞类型。为此,聚类已成为一项关键的计算技术,用于根据转录组特征对细胞进行分组,从而能够从每个细胞簇中识别后续的细胞类型。由于转录组的特征维度很高(即每个细胞中有大量的测量基因),并且只有一小部分基因是细胞类型特定的,因此对于生成细胞类型特定的簇来说,直接在原始特征/基因维度上进行聚类可能会导致没有信息的簇,从而阻碍正确的细胞类型识别。结果:本文提出了一种基于自动编码器的簇集成框架,该框架首先从数据中获取随机子空间投影,然后使用自动编码器人工神经网络将每个随机投影压缩到低维空间,最后在所有编码的数据集上应用集成聚类来生成细胞簇。我们使用四个评估指标来评估聚类性能,我们的实验表明,当同时应用标准k-Means聚类算法和专门为scRNA-seq数据设计的最先进的基于核的聚类算法(SIMLR)时,所提出的基于自动编码器的聚类集成可以显著改善特定细胞类型的聚类。与直接在原始数据集上使用这些聚类算法相比,根据所使用的评估指标,某些情况下的性能提升高达100%。结论:我们的结果表明,所提出的框架可以帮助更准确地识别细胞类型以及其他下游分析。创建所提出的基于自动编码器的集群集成框架的代码可从https://github.com/gedcom/scCCESS免费获得
Background: Single-cell RNA-sequencing (scRNA-seq) is a transformative technology, allowing global transcriptomes of individual cells to be profiled with high accuracy. An essential task in scRNA-seq data analysis is the identification of cell types from complex samples or tissues profiled in an experiment. To this end, clustering has become a key computational technique for grouping cells based on their transcriptome profiles, enabling subsequent cell type identification from each cluster of cells. Due to the high feature-dimensionality of the transcriptome (i.e. the large number of measured genes in each cell) and because only a small fraction of genes are cell type-specific and therefore informative for generating cell type-specific clusters, clustering directly on the original feature/gene dimension may lead to uninformative clusters and hinder correct cell type identification.Results: Here, we propose an autoencoder-based cluster ensemble framework in which we first take random subspace projections from the data, then compress each random projection to a low-dimensional space using an autoencoder artificial neural network, and finally apply ensemble clustering across all encoded datasets to generate clusters of cells. We employ four evaluation metrics to benchmark clustering performance and our experiments demonstrate that the proposed autoencoder-based cluster ensemble can lead to substantially improved cell type-specific clusters when applied with both the standard k-means clustering algorithm and a state-of-the-art kernel-based clustering algorithm (SIMLR) designed specifically for scRNA-seq data. Compared to directly using these clustering algorithms on the original datasets, the performance improvement in some cases is up to 100%, depending on the evaluation metric used.Conclusions: Our results suggest that the proposed framework can facilitate more accurate cell type identification as well as other downstream analyses. The code for creating the proposed autoencoder-based cluster ensemble framework is freely available from https://github.com/gedcom/scCCESS