RNA-Seq optimization with eQTL gold standards.

RNA-Seq optimization with eQTL gold standards.
复制标题

DOI:
10.1186/1471-2164-14-892
复制
发表时间:
2013-12-17
期刊:
影响因子:
4.4
通讯作者:
Arking DE
Arking DE
中科院分区:
生物学2区
文献类型:
--
作者:
Ellis SE;Gupta S;Ashar FN;Bader JS;West AB;Arking DE

文献摘要

参考文献

被引文献

相似文献

RNA测序(RNA-Seq)实验已经被优化用于文库制备、作图和基因表达估计。然而,这些方法在差异表达的下一阶段分析中暴露出弱点,结果对系统样本分层敏感,或者在更极端的情况下,对离群值敏感。此外,还缺乏一种评估对数据采取的标准化和调整措施的方法。为了解决这些问题,我们利用先前发表的eQTL作为一个新的黄金标准,在一个框架的中心,整合DNA基因型和RNA-Seq数据,以优化分析,并帮助理解遗传变异和基因表达。在检测样本污染和测序RNA-Seq数据中的离群值后,使用一组先前发表的脑eQTL来确定样本离群值去除是否合适。已知eQTL的复制改善支持在下游分析中去除这些样品。eQTL复制进一步用于评估标准化方法、协变量纳入和基因注释。该方法在来自GTEx项目的独立RNA-Seq血液数据集和组织适当的eQTL集中进行了验证。两个数据集中的eQTL复制突出了在RNA-Seq数据分析中考虑未知协变量的必要性。由于每个RNA-Seq实验都是独特的,具有其自身的实验特定限制,因此我们提供了一种易于实施的方法,该方法使用已知eQTL的复制来指导数据分析管道中的每个步骤。在本文提供的两个数据集中,我们不仅强调了仔细检测离群值的必要性,还强调了在RNA-Seq实验中考虑未知协变量的必要性。
RNA-Sequencing (RNA-Seq) experiments have been optimized for library preparation, mapping, and gene expression estimation. These methods, however, have revealed weaknesses in the next stages of analysis of differential expression, with results sensitive to systematic sample stratification or, in more extreme cases, to outliers. Further, a method to assess normalization and adjustment measures imposed on the data is lacking. To address these issues, we utilize previously published eQTLs as a novel gold standard at the center of a framework that integrates DNA genotypes and RNA-Seq data to optimize analysis and aid in the understanding of genetic variation and gene expression. After detecting sample contamination and sequencing outliers in RNA-Seq data, a set of previously published brain eQTLs was used to determine if sample outlier removal was appropriate. Improved replication of known eQTLs supported removal of these samples in downstream analyses. eQTL replication was further employed to assess normalization methods, covariate inclusion, and gene annotation. This method was validated in an independent RNA-Seq blood data set from the GTEx project and a tissue-appropriate set of eQTLs. eQTL replication in both data sets highlights the necessity of accounting for unknown covariates in RNA-Seq data analysis. As each RNA-Seq experiment is unique with its own experiment-specific limitations, we offer an easily-implementable method that uses the replication of known eQTLs to guide each step in one’s data analysis pipeline. In the two data sets presented herein, we highlight not only the necessity of careful outlier detection but also the need to account for unknown covariates in RNA-Seq experiments.
DOI: 10.1093/biostatistics/kxr054
发表时间: 2012-04
期刊: Biostatistics (Oxford, England)
影响因子: --
作者:
Hansen KD;Irizarry RA;Wu Z
通讯作者: Wu Z
DOI: 10.1371/journal.pcbi.1000770
发表时间: 2010-05-06
影响因子: 4.3
作者:
Stegle O;Parts L;Durbin R;Winn J
通讯作者: Winn J
DOI: 10.1186/1471-2105-12-480
发表时间: 2011-12-17
期刊: BMC bioinformatics
影响因子: 3
作者:
Risso D;Schwartz K;Sherlock G;Dudoit S
通讯作者: Dudoit S
通过替代变量分析捕获基因表达研究中的异质性。
DOI: 10.1371/journal.pgen.0030161
发表时间: 2007-09
期刊: PLOS GENETICS
影响因子: 4.5
作者:
Leek, Jeffrey T.;Storey, John D.
通讯作者: Storey, John D.
DOI: 10.1371/journal.pone.0068141
发表时间: 2013
期刊: PloS one
影响因子: 3.7
作者:
Mostafavi S;Battle A;Zhu X;Urban AE;Levinson D;Montgomery SB;Koller D
通讯作者: Koller D