Finding consistent patterns: a nonparametric approach for identifying differential expression in RNA-Seq data.

Finding consistent patterns: a nonparametric approach for identifying differential expression in RNA-Seq data.
复制标题

DOI:
10.1177/0962280211428386
复制
发表时间:
2013-10
影响因子:
2.3
通讯作者:
Tibshirani R
Tibshirani R
中科院分区:
医学3区
文献类型:
--
作者:
Li J;Tibshirani R

文献摘要

被引文献

相似文献

我们讨论了与RNA测序(RNA-Seq)和其他基于测序的比较基因组实验结果相关的特征的识别。RNA-Seq数据采用计数的形式,因此基于正态分布的模型通常不适用。这个问题特别具有挑战性,因为不同的测序实验可能会产生完全不同的读数总数或“测序深度”。现有的方法,这个问题是基于泊松或负二项模型:它们是有用的,但可以严重影响的数据中的“离群值”。我们介绍了一个简单的,非参数的方法与restaurant占不同的测序深度。新方法比参数方法更稳健。它可以应用于定量、生存、两类或多类结局的数据。我们比较我们提出的方法泊松和负二项式为基础的方法在模拟和真实的数据集,发现我们的方法发现更多的一致性模式比竞争的方法。
We discuss the identification of features that are associated with an outcome in RNA-Sequencing (RNA-Seq) and other sequencing-based comparative genomic experiments. RNA-Seq data takes the form of counts, so models based on the normal distribution are generally unsuitable. The problem is especially challenging because different sequencing experiments may generate quite different total numbers of reads, or ‘sequencing depths’. Existing methods for this problem are based on Poisson or negative binomial models: they are useful but can be heavily influenced by ‘outliers’ in the data. We introduce a simple, nonparametric method with resampling to account for the different sequencing depths. The new method is more robust than parametric methods. It can be applied to data with quantitative, survival, two-class or multiple-class outcomes. We compare our proposed method to Poisson and negative binomial-based methods in simulated and real data sets, and find that our method discovers more consistent patterns than competing methods.