Normalizing RNA-sequencing data by modeling hidden covariates with prior knowledge.

Normalizing RNA-sequencing data by modeling hidden covariates with prior knowledge.
复制标题

DOI:
10.1371/journal.pone.0068141
复制
发表时间:
2013
期刊:
影响因子:
3.7
通讯作者:
Koller D
Koller D
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Mostafavi S;Battle A;Zhu X;Urban AE;Levinson D;Montgomery SB;Koller D

文献摘要

参考文献

被引文献

相似文献

测量表达水平的转录组学测定广泛用于研究细胞过程中环境或遗传变异的表现。特别是RNA测序有可能大大提高这种理解,因为它能够分析整个转录组,包括新的转录事件。然而,与早期的表达测定一样,RNA测序数据的分析需要仔细考虑可能在表达测量中引入系统性、混杂变异性的因素,从而导致虚假相关性。在这里,我们考虑的问题建模和删除已知的和隐藏的混杂因素的影响,从RNA测序数据。我们描述了一个统一的残差框架,封装现有的方法,并使用这个框架,提出了一种新的方法,HCP(隐藏协变量与先验)。HCP使用关于混杂因素的更明智的假设,并且在具有低得多的计算成本的同时表现得与现有方法一样好或更好。我们的实验表明,考虑到已知和隐藏的因素与适当的模型提高了RNA测序数据的质量在两个非常不同的任务:检测与附近的表达变异(顺式eQTL),并构建准确的共表达网络相关的遗传变异。
Transcriptomic assays that measure expression levels are widely used to study the manifestation of environmental or genetic variations in cellular processes. RNA-sequencing in particular has the potential to considerably improve such understanding because of its capacity to assay the entire transcriptome, including novel transcriptional events. However, as with earlier expression assays, analysis of RNA-sequencing data requires carefully accounting for factors that may introduce systematic, confounding variability in the expression measurements, resulting in spurious correlations. Here, we consider the problem of modeling and removing the effects of known and hidden confounding factors from RNA-sequencing data. We describe a unified residual framework that encapsulates existing approaches, and using this framework, present a novel method, HCP (Hidden Covariates with Prior). HCP uses a more informed assumption about the confounding factors, and performs as well or better than existing approaches while having a much lower computational cost. Our experiments demonstrate that accounting for known and hidden factors with appropriate models improves the quality of RNA-sequencing data in two very different tasks: detecting genetic variations that are associated with nearby expression variations (cis-eQTLs), and constructing accurate co-expression networks.
DOI: 10.1186/gb-2008-9-s1-s2
发表时间: 2008
期刊: Genome biology
影响因子: 12.3
作者:
Peña-Castillo L;Tasan M;Myers CL;Lee H;Joshi T;Zhang C;Guan Y;Leone M;Pagnani A;Kim WK;Krumpelman C;Tian W;Obozinski G;Qi Y;Mostafavi S;Lin GN;Berriz GF;Gibbons FD;Lanckriet G;Qiu J;Grant C;Barutcuoglu Z;Hill DP;Warde-Farley D;Grouios C;Ray D;Blake JA;Deng M;Jordan MI;Noble WS;Morris Q;Klein-Seetharaman J;Bar-Joseph Z;Chen T;Sun F;Troyanskaya OG;Marcotte EM;Xu D;Hughes TR;Roth FP
通讯作者: Roth FP
DOI: 10.1371/journal.pcbi.1000770
发表时间: 2010-05-06
影响因子: 4.3
作者:
Stegle O;Parts L;Durbin R;Winn J
通讯作者: Winn J
DOI: 10.1186/gb-2010-11-3-r25
发表时间: 2010
期刊: Genome biology
影响因子: 12.3
作者:
Robinson MD;Oshlack A
通讯作者: Oshlack A
DOI: 10.1038/47048
发表时间: 1999-11-04
期刊: NATURE
影响因子: 64.8
作者:
Marcotte, EM;Pellegrini, M;Eisenberg, D
通讯作者: Eisenberg, D
DOI: 10.1371/journal.pgen.1001117
发表时间: 2010-09-16
期刊: PLoS genetics
影响因子: 4.5
作者:
Engelhardt BE;Stephens M
通讯作者: Stephens M