Random forest based similarity learning for single cell RNA sequencing data.

Random forest based similarity learning for single cell RNA sequencing data.
复制标题

DOI:
10.1093/bioinformatics/bty260
复制
发表时间:
2018-07-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Kostka D
Kostka D
中科院分区:
其他
文献类型:
--
作者:
Pouyan MB;Kostka D

文献摘要

参考文献

被引文献

相似文献

应用于单细胞的全基因组转录组测序(scRNA-seq)正迅速成为生物学和生物医学研究许多领域的一种选择。科学目标往往围绕着细胞类型或亚型的发现或表征,因此,从scRNA-seq数据中获得准确的细胞-细胞相似性是许多研究的关键步骤。虽然scRNA-seq数据分析工具的开发取得了快速进展,但明确解决这一任务的方法很少。此外,scRNA-seq数据集中存在的噪声的丰度和类型表明,通用方法或为大量RNA-seq数据开发的方法的应用可能不是最佳的。在这里,我们提出了RAFSIL,一种基于随机森林的方法,从scRNA-seq数据中学习细胞-细胞相似性。RAFSIL实现了一个两步的过程,其中针对scRNA-seq数据的特征构建之后是相似性学习。它被设计为具有适应性和可扩展性,并且RAFSIL相似性可以用于典型的探索性数据分析任务,如降维、可视化和聚类。我们表明,我们的方法在不同的数据集上与当前的方法相比具有优势,并且在其他方法失败的情况下,它可用于检测和突出scRNA-seq数据集中不需要的技术变化。总的来说,RAFSIL实现了一种灵活的方法,产生了一种有用的工具,可以改进scRNA-seq数据的分析。RAFSIL R软件包可在www.kostkalab.net/software.html上获得,补充数据可在Bioinformatics在线获得。
Genome-wide transcriptome sequencing applied to single cells (scRNA-seq) is rapidly becoming an assay of choice across many fields of biological and biomedical research. Scientific objectives often revolve around discovery or characterization of types or sub-types of cells, and therefore, obtaining accurate cell–cell similarities from scRNA-seq data is a critical step in many studies. While rapid advances are being made in the development of tools for scRNA-seq data analysis, few approaches exist that explicitly address this task. Furthermore, abundance and type of noise present in scRNA-seq datasets suggest that application of generic methods, or of methods developed for bulk RNA-seq data, is likely suboptimal. Here, we present RAFSIL, a random forest based approach to learn cell–cell similarities from scRNA-seq data. RAFSIL implements a two-step procedure, where feature construction geared towards scRNA-seq data is followed by similarity learning. It is designed to be adaptable and expandable, and RAFSIL similarities can be used for typical exploratory data analysis tasks like dimension reduction, visualization and clustering. We show that our approach compares favorably with current methods across a diverse collection of datasets, and that it can be used to detect and highlight unwanted technical variation in scRNA-seq datasets in situations where other methods fail. Overall, RAFSIL implements a flexible approach yielding a useful tool that improves the analysis of scRNA-seq data. The RAFSIL R package is available at www.kostkalab.net/software.html Supplementary data are available at Bioinformatics online.
DOI: 10.1016/j.cell.2016.01.047
发表时间: 2016-03-24
期刊: Cell
影响因子: 64.5
作者:
Goolam M;Scialdone A;Graham SJL;Macaulay IC;Jedrusik A;Hupalowska A;Voet T;Marioni JC;Zernicka-Goetz M
通讯作者: Zernicka-Goetz M
DOI: 10.1038/nmeth.4236
发表时间: 2017-05-01
期刊: NATURE METHODS
影响因子: 48
作者:
Kiselev, Vladimir Yu;Kirschner, Kristina;Hemberg, Martin
通讯作者: Hemberg, Martin
DOI: 10.1093/bioinformatics/bth294
发表时间: 2004-11-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Lanckriet, GRG;De Bie, T;Noble, WS
通讯作者: Noble, WS
DOI: 10.1186/s13059-016-0881-8
发表时间: 2016-01-26
期刊: Genome biology
影响因子: 12.3
作者:
Conesa A;Madrigal P;Tarazona S;Gomez-Cabrero D;Cervera A;McPherson A;Szcześniak MW;Gaffney DJ;Elo LL;Zhang X;Mortazavi A
通讯作者: Mortazavi A
DOI: 10.1038/nature14966
发表时间: 2015-09-10
期刊: NATURE
影响因子: 64.8
作者:
Grun, Dominic;Lyubimova, Anna;van Oudenaarden, Alexander
通讯作者: van Oudenaarden, Alexander