Normalization, testing, and false discovery rate estimation for RNA-sequencing data

Normalization, testing, and false discovery rate estimation for RNA-sequencing data
复制标题

DOI:
10.1093/biostatistics/kxr031
复制
发表时间:
2012-07-01
期刊:
影响因子:
2.1
通讯作者:
Tibshirani, Robert
Tibshirani, Robert
中科院分区:
数学2区
文献类型:
--
作者:
Li, Jun;Witten, Daniela M.;Tibshirani, Robert

文献摘要

被引文献

相似文献

我们讨论了与RNA测序和其他基于序列的比较基因组实验的结果相关的基因的识别。RNA测序数据采用计数的形式,因此基于高斯分布的模型是不合适的。此外,标准化是具有挑战性的,因为不同的测序实验可能产生完全不同的读数总数。为了克服这些困难,我们使用对数线性模型与一种新的方法来规范化。我们推导出一种新的方法来估计错误发现率(FDR)。我们的方法可以应用于定量,两类或多类结果的数据,即使是大数据集的计算速度也很快。我们研究了我们的方法的显着性计算和FDR估计的准确性,我们证明了我们的方法具有潜在的优势,现有的方法是基于泊松或负二项模型。总之,这项工作为测序数据的显著性分析提供了一个管道。
We discuss the identification of genes that are associated with an outcome in RNA sequencing and other sequence-based comparative genomic experiments. RNA-sequencing data take the form of counts, so models based on the Gaussian distribution are unsuitable. Moreover, normalization is challenging because different sequencing experiments may generate quite different total numbers of reads. To overcome these difficulties, we use a log-linear model with a new approach to normalization. We derive a novel procedure to estimate the false discovery rate (FDR). Our method can be applied to data with quantitative, two-class, or multiple-class outcomes, and the computation is fast even for large data sets. We study the accuracy of our approaches for significance calculation and FDR estimation, and we demonstrate that our method has potential advantages over existing methods that are based on a Poisson or negative binomial model. In summary, this work provides a pipeline for the significance analysis of sequencing data.