A permutation-based non-parametric analysis of CRISPR screen data.

A permutation-based non-parametric analysis of CRISPR screen data.
复制标题

DOI:
10.1186/s12864-017-3938-5
复制
发表时间:
2017-07-19
期刊:
影响因子:
4.4
通讯作者:
Xiao G
Xiao G
中科院分区:
生物学2区
文献类型:
--
作者:
Jia G;Wang X;Xiao G

文献摘要

参考文献

被引文献

相似文献

通常在培养的细胞中实施规则间隔短回文重复序列(CRISPR)筛选以鉴定具有关键功能的基因。尽管已经开发或采用了几种方法来分析CRISPR筛选数据,但没有一种特定的算法得到普及。因此,需要严格的程序来克服现有算法的缺点。我们开发了一种基于排列的非参数分析(PBNPA)算法,该算法通过排列sgRNA标签来计算基因水平的p值,从而避免了限制性分布假设。虽然PBNPA被设计用于分析CRISPR数据,但它也可以应用于分析用siRNA或shRNA实施的遗传筛选和药物筛选。在模拟数据和真实的数据上,我们比较了PBNPA与竞争方法的性能。PBNPA在各种设置下模拟数据的受试者操作特征(ROC)曲线和错误发现率(FDR)控制方面优于最近设计用于CRISPR筛选分析的方法,以及用于分析其他功能基因组学筛选的方法。值得注意的是,PBNPA算法在已发布的真实的数据上也表现出更好的一致性和FDR控制。PBNPA比其竞争对手产生更一致和可靠的结果,特别是当数据质量较低时。PBNPA的R包可在https://cran.r-project.org/web/packages/PBNPA/上获得。本文的在线版本(doi:10.1186/s12864-017-3938-5)包含补充材料,可供授权用户使用。
Clustered regularly-interspaced short palindromic repeats (CRISPR) screens are usually implemented in cultured cells to identify genes with critical functions. Although several methods have been developed or adapted to analyze CRISPR screening data, no single specific algorithm has gained popularity. Thus, rigorous procedures are needed to overcome the shortcomings of existing algorithms. We developed a Permutation-Based Non-Parametric Analysis (PBNPA) algorithm, which computes p-values at the gene level by permuting sgRNA labels, and thus it avoids restrictive distributional assumptions. Although PBNPA is designed to analyze CRISPR data, it can also be applied to analyze genetic screens implemented with siRNAs or shRNAs and drug screens. We compared the performance of PBNPA with competing methods on simulated data as well as on real data. PBNPA outperformed recent methods designed for CRISPR screen analysis, as well as methods used for analyzing other functional genomics screens, in terms of Receiver Operating Characteristics (ROC) curves and False Discovery Rate (FDR) control for simulated data under various settings. Remarkably, the PBNPA algorithm showed better consistency and FDR control on published real data as well. PBNPA yields more consistent and reliable results than its competitors, especially when the data quality is low. R package of PBNPA is available at: https://cran.r-project.org/web/packages/PBNPA/. The online version of this article (doi:10.1186/s12864-017-3938-5) contains supplementary material, which is available to authorized users.
DOI: 10.1101/gr.162339.113
发表时间: 2014-01
期刊: Genome research
影响因子: 7
作者:
Cho SW;Kim S;Kim Y;Kweon J;Kim HS;Bae S;Kim JS
通讯作者: Kim JS
DOI: 10.1093/bib/bbt086
发表时间: 2015-01
影响因子: 9.5
作者:
Seyednasrollah F;Laiho A;Elo LL
通讯作者: Elo LL
DOI: 10.1093/bioinformatics/bti448
发表时间: 2005-07-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Pawitan, Y;Michiels, S;Ploner, A
通讯作者: Ploner, A
DOI: 10.1214/12-aoas592
发表时间: 2013-03-01
期刊: The annals of applied statistics
影响因子: --
作者:
Chen J;Li H
通讯作者: Li H
DOI: 10.1038/nbt.3567
发表时间: 2016-06
影响因子: 46.9
作者:
Morgens DW;Deans RM;Li A;Bassik MC
通讯作者: Bassik MC