Statistical models for RNA-seq data derived from a two-condition 48-replicate experiment.

Statistical models for RNA-seq data derived from a two-condition 48-replicate experiment.
复制标题

DOI:
10.1093/bioinformatics/btv425
复制
发表时间:
2015-11-15
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Barton GJ
Barton GJ
中科院分区:
其他
文献类型:
--
作者:
Gierliński M;Cole C;Schofield P;Schurch NJ;Sherstnev A;Singh V;Wrobel N;Gharbi K;Simpson G;Owen-Hughes T;Blaxter M;Barton GJ

文献摘要

被引文献

相似文献

动机:高通量 RNA 测序 (RNA-seq) 现在是确定差异基因表达的标准方法。识别差异表达基因关键取决于对读数计数变异性的估计。这些估计通常基于统计模型,例如 EdgeR、DESeq 和 cuffdiff 工具所采用的负二项分布。到目前为止,这些模型的有效性通常是在低重复 RNA-seq 数据或模拟上进行测试的。结果:在酵母中进行了 48 次重复的 RNA-seq 实验,并根据理论模型测试了数据。观察到的基因读数计数符合对数正态分布和负二项分布,而均值-方差关系遵循~0.01的恒定离散参数线。高重复数据还可以进行严格的质量控制和筛选“不良”重复,这可能会极大地影响基因读数的分布。可用性和实施​​:RNA-seq 数据已提交给 ENA 档案,项目 ID PRJEB5348。联系方式:g.j.barton@dundee.ac.uk
Motivation: High-throughput RNA sequencing (RNA-seq) is now the standard method to determine differential gene expression. Identifying differentially expressed genes crucially depends on estimates of read-count variability. These estimates are typically based on statistical models such as the negative binomial distribution, which is employed by the tools edgeR, DESeq and cuffdiff. Until now, the validity of these models has usually been tested on either low-replicate RNA-seq data or simulations. Results: A 48-replicate RNA-seq experiment in yeast was performed and data tested against theoretical models. The observed gene read counts were consistent with both log-normal and negative binomial distributions, while the mean-variance relation followed the line of constant dispersion parameter of ∼0.01. The high-replicate data also allowed for strict quality control and screening of ‘bad’ replicates, which can drastically affect the gene read-count distribution. Availability and implementation: RNA-seq data have been submitted to ENA archive with project ID PRJEB5348. Contact: g.j.barton@dundee.ac.uk