A multitask clustering approach for single-cell RNA-seq analysis in Recessive Dystrophic Epidermolysis Bullosa.

A multitask clustering approach for single-cell RNA-seq analysis in Recessive Dystrophic Epidermolysis Bullosa.
复制标题

DOI:
10.1371/journal.pcbi.1006053
复制
发表时间:
2018-04
影响因子:
4.3
通讯作者:
Tolar J
Tolar J
中科院分区:
生物学2区
文献类型:
--
作者:
Zhang H;Lee CAA;Li Z;Garbe JR;Eide CR;Petegrosso R;Kuang R;Tolar J

文献摘要

参考文献

被引文献

相似文献

单细胞 RNA 测序 (scRNA-seq) 已广泛应用于通过检测异质细胞群中的亚群来发现新的细胞类型。由于与批量 RNA-seq 实验相比,scRNA-seq 实验的读取覆盖率/标签计数较低,并引入了更多技术偏差,因此有限数量的采样细胞与实验偏差和其他数据集特定变异相结合,对跨数据集分析和跨多个细胞群的相关生物学变异的发现提出了挑战。在本文中,我们介绍了一种单细胞 RNA-seq 数据(scVDMC)的方差驱动多任务聚类方法,该方法利用来自生物重复或不同样本的多个单细胞群体。 scVDMC 在具有相似细胞类型和标记但不同表达模式的多个 scRNA-seq 实验中对单细胞进行聚类,因此 scRNA-seq 数据比仅增加样本量的典型汇总分析更好地整合。通过控制每个数据集中和所有数据集中细胞簇之间的差异,scVDMC 可以使用共享的细胞类型标记但在所有实验中使用不同的簇中心来检测每个单独实验中的细胞亚群。应用于两个具有多个重复的真实 scRNA-seq 数据集和三个患者样本的一个基于液滴的大规模数据集时,scVDMC 比池聚类和其他最近提出的 scRNA-seq 聚类方法更准确地检测细胞群和已知细胞标记。在应用于内部隐性营养不良性大疱性表皮松解症 (RDEB) scRNA-seq 数据的案例研究中,scVDMC 揭示了几种新的细胞类型和通过流式细胞术验证的未知标记。 MATLAB/Octave 代码可在 https://github.com/kuanglab/scVDMC 获取。 scRNA-seq 能够对异质细胞群进行详细分析,并可用于揭示谱系关系或发现新的细胞类型。在文献中,几乎没有致力于开发用于多个单细胞群体的跨群体转录组分析的计算方法。跨细胞群聚类问题不同于传统的聚类问题,因为单细胞群可以从不同的患者、不同的组织样本或不同的实验重复中收集。伴随的生物学和技术变化往往主导对来自多个群体的汇集的单细胞进行聚类的信号。在这项工作中,我们开发了一种多任务聚类方法来解决跨群体聚类问题。该方法同时对每个单独的细胞群进行聚类,并控制每个细胞群内和细胞群之间的细胞类型聚类中心之间的差异。我们证明,我们的多任务聚类方法显着提高了三个公共 scRNA-seq 数据集中的聚类准确性和标记发现,并将该方法应用于内部隐性营养不良性大疱性表皮松解症 (RDEB) 数据集。我们的结果表明,多任务聚类是 scRNA-seq 数据跨群体分析的一种有前景的新方法。
Single-cell RNA sequencing (scRNA-seq) has been widely applied to discover new cell types by detecting sub-populations in a heterogeneous group of cells. Since scRNA-seq experiments have lower read coverage/tag counts and introduce more technical biases compared to bulk RNA-seq experiments, the limited number of sampled cells combined with the experimental biases and other dataset specific variations presents a challenge to cross-dataset analysis and discovery of relevant biological variations across multiple cell populations. In this paper, we introduce a method of variance-driven multitask clustering of single-cell RNA-seq data (scVDMC) that utilizes multiple single-cell populations from biological replicates or different samples. scVDMC clusters single cells in multiple scRNA-seq experiments of similar cell types and markers but varying expression patterns such that the scRNA-seq data are better integrated than typical pooled analyses which only increase the sample size. By controlling the variance among the cell clusters within each dataset and across all the datasets, scVDMC detects cell sub-populations in each individual experiment with shared cell-type markers but varying cluster centers among all the experiments. Applied to two real scRNA-seq datasets with several replicates and one large-scale droplet-based dataset on three patient samples, scVDMC more accurately detected cell populations and known cell markers than pooled clustering and other recently proposed scRNA-seq clustering methods. In the case study applied to in-house Recessive Dystrophic Epidermolysis Bullosa (RDEB) scRNA-seq data, scVDMC revealed several new cell types and unknown markers validated by flow cytometry. MATLAB/Octave code available at https://github.com/kuanglab/scVDMC. scRNA-seq enables detailed profiling of heterogeneous cell populations and can be used to reveal lineage relationships or discover new cell types. In the literature, there has been little effort directed towards developing computational methods for cross-population transcriptome analysis of multiple single-cell populations. The cross-cell-population clustering problem is different from the traditional clustering problem because single-cell populations can be collected from different patients, different samples of a tissue, or different experimental replicates. The accompanying biological and technical variation tends to dominate the signals for clustering the pooled single cells from the multiple populations. In this work, we have developed a multitask clustering method to address the cross-population clustering problem. The method simultaneously clusters each individual cell population and controls variance among the cell-type cluster centers within each cell population and across the cell populations. We demonstrate that our multitask clustering method significantly improves clustering accuracy and marker discovery in three public scRNA-seq datasets and also apply the method to an in-house Recessive Dystrophic Epidermolysis Bullosa (RDEB) dataset. Our results make it evident that multitask clustering is a promising new approach for cross-population analysis of scRNA-seq data.
DOI: 10.1186/s13059-016-0927-y
发表时间: 2016-04-07
期刊: Genome biology
影响因子: 12.3
作者:
Bacher R;Kendziorski C
通讯作者: Kendziorski C
DOI: 10.1164/rccm.201509-1863oc
发表时间: 2016-05-15
影响因子: 24.7
作者:
Mathai, Susan K.;Pedersen, Brent S.;Schwartz, David A.
通讯作者: Schwartz, David A.
DOI: 10.3390/biology1030658
发表时间: 2012-11-16
期刊: Biology
影响因子: 4.2
作者:
Hebenstreit D
通讯作者: Hebenstreit D
DOI: 10.1083/jcb.104.3.611
发表时间: 1987-03-01
影响因子: 7.8
作者:
KEENE, DR;SAKAI, LY;BURGESON, RE
通讯作者: BURGESON, RE
DOI: 10.1038/nmeth.4236
发表时间: 2017-05-01
期刊: NATURE METHODS
影响因子: 48
作者:
Kiselev, Vladimir Yu;Kirschner, Kristina;Hemberg, Martin
通讯作者: Hemberg, Martin