Correction for hidden confounders in the genetic analysis of gene expression

Correction for hidden confounders in the genetic analysis of gene expression
复制标题

DOI:
10.1073/pnas.1002425107
复制
发表时间:
2010-09-21
影响因子:
11.1
通讯作者:
Heckerman, David
Heckerman, David
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Listgarten, Jennifer;Kadie, Carl;Heckerman, David

文献摘要

被引文献

相似文献

了解疾病的遗传基础对于筛查、治疗、药物开发和基本生物学洞察非常重要。获得这种理解的一种方法是找出 DNA 的哪些部分(例如单核苷酸多态性)影响特定的中间过程(例如基因表达)。天真的,可以通过对遗传变异和基因转录本的所有配对组合进行简单的统计测试来识别这种关联。然而,数据中隐藏着各种各样的混杂因素,如果处理不当,会导致虚假关联和遗漏关联。我们提出了一种统计模型,当这些混杂因素未知时,可以联合纠正两种特定类型的隐藏结构-群体结构(例如种族、家庭相关性)和微阵列表达伪影(例如批次效应)。将我们的方法应用于真实和合成、人类和小鼠数据,我们证明了对混杂因素进行联合校正的必要性,以及基于当前文献中的其他可能方法的缺点。特别是,我们表明我们的模型类别具有在合成数据上检测 eQTL 的最大能力,并且在应用于实际数据的青铜标准上具有最佳性能。最后,我们的软件以及我们发现的与它的关联可以在 http://www.microsoft.com/science 上找到。
Understanding the genetic underpinnings of disease is important for screening, treatment, drug development, and basic biological insight. One way of getting at such an understanding is to find out which parts of our DNA, such as single-nucleotide polymorphisms, affect particular intermediary processes such as gene expression. Naively, such associations can be identified using a simple statistical test on all paired combinations of genetic variants and gene transcripts. However, a wide variety of confounders lie hidden in the data, leading to both spurious associations and missed associations if not properly addressed. We present a statistical model that jointly corrects for two particular kinds of hidden structure-population structure (e.g., race, family-relatedness), and microarray expression artifacts (e.g., batch effects), when these confounders are unknown. Applying our method to both real and synthetic, human and mouse data, we demonstrate the need for such a joint correction of confounders, and also the disadvantages of other possible approaches based on those in the current literature. In particular, we show that our class of models has maximum power to detect eQTL on synthetic data, and has the best performance on a bronze standard applied to real data. Lastly, our software and the associations we found with it are available at http://www.microsoft.com/science.