Iterative Usage of Fixed and Random Effect Models for Powerful and Efficient Genome-Wide Association Studies.

Iterative Usage of Fixed and Random Effect Models for Powerful and Efficient Genome-Wide Association Studies.
复制标题

DOI:
10.1371/journal.pgen.1005767
复制
发表时间:
2016-02
期刊:
影响因子:
4.5
通讯作者:
Zhang Z
Zhang Z
中科院分区:
生物学2区
文献类型:
--
作者:
Liu X;Huang M;Fan B;Buckler ES;Zhang Z

文献摘要

被引文献

相似文献

全基因组关联研究(GWAS)中的假阳性可以通过固定效应和随机效应混合线性模型(MLM)来有效控制,该模型将个体之间的群体结构和亲属关系纳入其中,以调整标记物的关联检验;然而,调整也会损害真阳性。改进的MLM方法,多位点线性混合模型(MLMM),采用多个标记同时作为协变量的逐步MLM,部分消除测试标记和亲属关系之间的混杂。为了完全消除混杂,我们将MLMM分为两个部分:固定效应模型(FEM)和随机效应模型(REM),并迭代使用它们。FEM包含测试标志物(一次一个)和多个相关标志物作为协变量以控制假阳性。为了避免有限元中的模型过拟合问题,在REM中估计相关标记,使用它们来定义亲属关系。在每次迭代中,测试标记物和相关标记物的P值是统一的。我们将新方法命名为固定和随机模型循环概率统一(FarmCPU)。真实的和模拟的数据分析表明,FarmCPU提高了统计能力相比,目前的方法。额外的益处包括与个体数量和标记数量都呈线性的有效计算时间。现在,一个包含50万个个体和50万个标记的数据集可以在三天内进行分析。全基因组关联研究(GWAS)可以揭示遗传-表型关系,但有局限性。为了控制假阳性,人口结构和亲属关系被纳入一个固定和随机效应的混合线性模型(MLM)。然而,由于种群结构、亲缘关系和数量性状核苷酸(QTN)之间的混杂,MLM导致假阴性,错过了一些潜在的重要发现。本文提出了一种新的方法--固定与随机模型循环概率统一法(FarmCPU)。FarmCPU在固定效应模型中以相关标记作为协变量进行标记检验,并在随机效应模型中对相关协变量标记进行优化。该过程实现了有效的计算,消除了混淆,防止了模型过度拟合,同时控制了误报。FarmCPU控制误报以及MLM,减少误报和计算时间。研究人员不仅能够分析大数据,而且在绘制感兴趣的基因图谱时,还将取得更大的成功,错误更少。
False positives in a Genome-Wide Association Study (GWAS) can be effectively controlled by a fixed effect and random effect Mixed Linear Model (MLM) that incorporates population structure and kinship among individuals to adjust association tests on markers; however, the adjustment also compromises true positives. The modified MLM method, Multiple Loci Linear Mixed Model (MLMM), incorporates multiple markers simultaneously as covariates in a stepwise MLM to partially remove the confounding between testing markers and kinship. To completely eliminate the confounding, we divided MLMM into two parts: Fixed Effect Model (FEM) and a Random Effect Model (REM) and use them iteratively. FEM contains testing markers, one at a time, and multiple associated markers as covariates to control false positives. To avoid model over-fitting problem in FEM, the associated markers are estimated in REM by using them to define kinship. The P values of testing markers and the associated markers are unified at each iteration. We named the new method as Fixed and random model Circulating Probability Unification (FarmCPU). Both real and simulated data analyses demonstrated that FarmCPU improves statistical power compared to current methods. Additional benefits include an efficient computing time that is linear to both number of individuals and number of markers. Now, a dataset with half million individuals and half million markers can be analyzed within three days. Genome-Wide Association Studies (GWAS) can reveal genetic-phenotypic relationships, but have limitations. To control false positives, population structure and kinship are incorporated in a fixed and random effect Mixed Linear Model (MLM). However, because of the confounding between population structure, kinship, and quantitative trait nucleotides (QTNs), MLM leads to false negatives, missing some potentially important discoveries. Here, we present a new method, Fixed and random model Circulating Probability Unification (FarmCPU). FarmCPU performs marker tests with associated markers as covariates in a fixed effect model and optimization on the associated covariate markers in a random effect model separately. This process enables efficient computation, removes the confounding, prevents model over-fitting, and controls false positives simultaneously. FarmCPU controls false positives as well as MLM with reductions in both false negatives and computing times. Researchers will not only be able to analyze big data, but will also have greater success with fewer mistakes when mapping genes of interest.