Correcting gene expression data when neither the unwanted variation nor the factor of interest are observed.

Correcting gene expression data when neither the unwanted variation nor the factor of interest are observed.
复制标题

DOI:
10.1093/biostatistics/kxv026
复制
发表时间:
2016-01
期刊:
Biostatistics (Oxford, England)
影响因子:
--
通讯作者:
Speed TP
Speed TP
中科院分区:
其他
文献类型:
--
作者:
Jacob L;Gagnon-Bartsch JA;Speed TP

文献摘要

被引文献

相似文献

在处理大规模基因表达研究时,观察结果通常会受到平台或批次等不需要的变异来源的污染。在分析数据时不考虑这种不必要的变化可能会导致虚假的关联和丢失重要信号。当分析是无监督的,例如,当目标是聚类样本或建立一个修正版本的样本集时,而不是研究一个观察到的感兴趣的因素,考虑不必要的变化可能成为一项艰巨的任务。驱动不希望的变化的因素可能与未观察到的感兴趣的因素相关,因此,如果不小心进行,对前者的校正可能会删除后者。我们展示了如何阴性对照基因和重复样本可以用来估计不必要的基因表达的变化,并讨论如何使用这些信息来纠正表达数据。所提出的方法,然后评估合成数据和三个基因表达数据集。它们通常能够在不丢失感兴趣的信号的情况下消除不需要的变化,并且与最先进的校正相比毫不逊色。所有提出的方法都在生物导体包RUVnormalize中实现。
When dealing with large scale gene expression studies, observations are commonly contaminated by sources of unwanted variation such as platforms or batches. Not taking this unwanted variation into account when analyzing the data can lead to spurious associations and to missing important signals. When the analysis is unsupervised, e.g. when the goal is to cluster the samples or to build a corrected version of the dataset—as opposed to the study of an observed factor of interest—taking unwanted variation into account can become a difficult task. The factors driving unwanted variation may be correlated with the unobserved factor of interest, so that correcting for the former can remove the latter if not done carefully. We show how negative control genes and replicate samples can be used to estimate unwanted variation in gene expression, and discuss how this information can be used to correct the expression data. The proposed methods are then evaluated on synthetic data and three gene expression datasets. They generally manage to remove unwanted variation without losing the signal of interest and compare favorably to state-of-the-art corrections. All proposed methods are implemented in the bioconductor package RUVnormalize.