RgCop-A regularized copula based method for gene selection in single-cell RNA-seq data.

RgCop-A regularized copula based method for gene selection in single-cell RNA-seq data.
复制标题

DOI:
10.1371/journal.pcbi.1009464
复制
发表时间:
2021-10
影响因子:
4.3
通讯作者:
Bandyopadhyay S
Bandyopadhyay S
中科院分区:
生物学2区
文献类型:
--
作者:
Lall S;Ray S;Bandyopadhyay S

文献摘要

参考文献

被引文献

相似文献

在未注释的大型单细胞RNA测序(scRNA-seq)数据中的基因选择是下游分析的初始步骤中重要且关键的步骤。现有的方法主要基于高变异(高变异基因)或显着高表达(高表达基因),由于数据中存在技术噪音,无法提供稳定和预测的特征集。在这里,我们提出了RgCop,这是一种新的基于正则化copula的方法,用于从大型单细胞RNA-seq数据中进行基因选择。RgCop利用Copula相关性(Ccor),这是一种稳健的公平依赖性度量,其捕获单细胞表达数据中的一组基因之间的多变量依赖性。我们制定了一个目标函数,通过添加l1正则化项与Ccor惩罚的冗余系数的功能/基因,从而产生非冗余的有效功能/基因集。结果表明,与其他现有技术相比,RgCop在真实的scRNA-seq数据的聚类/分类性能上有显著提高。由于copula的尺度不变特性,RgCop在捕获噪声数据特征之间的依赖性方面表现得非常好,从而提高了方法的稳定性。此外,发现从scRNA-seq数据簇中识别的差异表达(DE)基因可以提供细胞的准确注释。最后,从RgCop获得的特征/基因能够以高精度注释未知细胞。基于高变异(高度可变基因)或显著高表达(高度表达基因)的基因选择的现有方法未能提供稳定和预测性特征/基因集。由于单细胞数据易受技术噪声的影响,因此在聚类之前选择的基因的质量在下游分析的初步步骤中至关重要。在这里,我们提出了一种新的正则化copula为基础的基因选择方法,利用copula相关性(Ccor)的措施,捕捉细胞间的变异性的数据。建议的目标函数使用l1正则化项来惩罚特征/基因的冗余系数。我们在细胞的聚类/分类性能方面比其他最先进的方法有了显着的改进。由于Copula的尺度不变特性,RgCop不受技术噪声的影响,这是scRNA-seq数据分析中的一个严重问题。此外,所选择的特征/基因能够以高精度确定未知细胞。最后,RgCop可适用于鉴定单细胞数据中的罕见细胞簇或次要亚群。
Gene selection in unannotated large single cell RNA sequencing (scRNA-seq) data is important and crucial step in the preliminary step of downstream analysis. The existing approaches are primarily based on high variation (highly variable genes) or significant high expression (highly expressed genes) failed to provide stable and predictive feature set due to technical noise present in the data. Here, we propose RgCop, a novel regularized copula based method for gene selection from large single cell RNA-seq data. RgCop utilizes copula correlation (Ccor), a robust equitable dependence measure that captures multivariate dependency among a set of genes in single cell expression data. We formulate an objective function by adding l1 regularization term with Ccor to penalizes the redundant co-efficient of features/genes, resulting non-redundant effective features/genes set. Results show a significant improvement in the clustering/classification performance of real life scRNA-seq data over the other state-of-the-art. RgCop performs extremely well in capturing dependence among the features of noisy data due to the scale invariant property of copula, thereby improving the stability of the method. Moreover, the differentially expressed (DE) genes identified from the clusters of scRNA-seq data are found to provide an accurate annotation of cells. Finally, the features/genes obtained from RgCop is able to annotate the unknown cells with high accuracy. The existing approaches for gene selection which are based on high variation (highly variable genes) or significant high expression (highly expressed genes), failed to provide a stable and predictive feature/gene set. Since single cell data is susceptible to technical noise, the quality of genes selected prior to clustering is of crucial importance in the preliminary steps of downstream analysis. Here, we propose a novel regularized copula based method for gene selection that leverage copula correlation (Ccor) measure for capturing cell-to-cell variability within the data. The proposed objective function uses an l1 regularization term to penalizes the redundant co-efficient of features/genes. We got significant improvement in the clustering/classification performance of cells over the other state-of-the-art. Due to the scale-invariant property of copula RgCop is impervious to technical noise, an acute issue associated with scRNA-seq data analysis. Moreover, the selected features/genes can be able to determine the unknown cells with high accuracy. Finally, RgCop can be applicable for identifying rare cell clusters or minor subpopulations within the single cell data.
DOI: 10.1016/j.cell.2018.07.028
发表时间: 2018-08-09
期刊: Cell
影响因子: 64.5
作者:
Saunders A;Macosko EZ;Wysoker A;Goldman M;Krienen FM;de Rivera H;Bien E;Baum M;Bortolin L;Wang S;Goeva A;Nemesh J;Kamitaki N;Brumbaugh S;Kulp D;McCarroll SA
通讯作者: McCarroll SA
DOI: 10.1038/nmeth.4236
发表时间: 2017-05-01
期刊: NATURE METHODS
影响因子: 48
作者:
Kiselev, Vladimir Yu;Kirschner, Kristina;Hemberg, Martin
通讯作者: Hemberg, Martin
DOI: 10.1186/1471-2105-9-225
发表时间: 2008-05-01
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Kim, Jong-Min;Jung, Yoon-Sung;Sohn, Insuk
通讯作者: Sohn, Insuk
DOI: 10.1038/s41467-018-07234-6
发表时间: 2018-11-09
影响因子: 16.6
作者:
Jindal A;Gupta P;Jayadeva;Sengupta D
通讯作者: Sengupta D
DOI: 10.1093/nar/gkw430
发表时间: 2016-07-27
影响因子: 14.9
作者:
Ji Z;Ji H
通讯作者: Ji H