Efficient control of population structure in model organism association mapping

Efficient control of population structure in model organism association mapping
复制标题

DOI:
10.1534/genetics.107.080101
复制
发表时间:
2008-03-01
期刊:
影响因子:
3.3
通讯作者:
Eskin, Eleazar
Eskin, Eleazar
中科院分区:
生物学2区
文献类型:
--
作者:
Kang, Hyun Min;Zaitlen, Noah A.;Eskin, Eleazar

文献摘要

被引文献

相似文献

在模式生物如近交系小鼠品系中进行全基因组关联作图是鉴定与人类疾病相关的风险因素的一种有前途的方法。然而,近交模式生物的遗传关联研究面临着菌株间复杂群体结构的问题。这会导致虚增的假阳性率,这无法使用人类关联研究中应用的标准方法(如基因组控制或结构化关联)进行校正。最近的研究表明,混合模型成功地纠正了玉米和拟南芥面板数据集的关联作图中的遗传相关性。然而,目前可用的混合模型方法遭受泡沫计算效率低下。在这篇文章中,我们提出了一种新的方法,有效的混合模型关联(EMMA),它纠正人口结构和遗传相关性的模式生物关联映射。我们的方法利用了优化问题的特定性质,将混合模型应用于关联捕捉,这使我们能够大幅提高计算速度和结果的可靠性。除了拟南芥和玉米数据集,我们还将EMMA应用于涉及数十万个SNP的近交系小鼠品系的计算机全基因组关联作图。我们还进行了广泛的模拟研究,以估计EMMA在各种SNP效应、不同程度的群体结构和每个菌株不同数量的多次测量下的统计功效。尽管由于可用的近交系数量有限,近交系小鼠关联作图的能力很强,但我们能够鉴定出显著相关的SNP,这些SNP落入通过先前研究鉴定的已知QTL或基因中,同时避免了假阳性的膨胀。我们的EMMA方法的R包实现和Web服务器是公开可用的。
Genomewide association mapping in model organisms such as inbred mouse strains is a promising approach for the identification of risk factors related to human diseases. However, genetic association studies in inbred model organisms are confronted by the problem of complex population structure among strains. This induces inflated false positive rates, which cannot be corrected using standard approaches applied in human association studies such as genomic control or structured association. Recent studies demonstrated that mixed models successfully correct for the genetic relatedness in association mapping in maize and Arabidopsis panel data sets. However, the currently available mixed-model methods suffer froth computational inefficiency. In this article, we propose a new method, efficient mixed-model association (EMMA), which corrects for population structure and genetic relatedness in model organism association mapping. Our method takes advantage of the specific nature of the optimization problem in applying mixed models for association snapping, which allows us to substantially increase the computational speed and reliability of the results. We applied EMMA to in silico whole-genome association mapping of inbred mouse strains involving hundreds of thousands of SNPs, in addition to Arabidopsis and maize data sets. We also performed extensive simulation studies to estimate the statistical power of EMMA under various SNP effects, varying degrees of population structure, and differing numbers of multiple measurements per strain. Despite the fruited power of inbred mouse association mapping due to the limited number of available inbred strains, we are able to identify significantly associated SNPs, which fall into known QTL or genes identified through previous studies while avoiding an inflation of false positives. An R package implementation and webserver of our EMMA method are publicly available.