Multimarker analysis and imputation of multiple platform pooling-based genome-wide association studies

Multimarker analysis and imputation of multiple platform pooling-based genome-wide association studies
复制标题

DOI:
10.1093/bioinformatics/btn333
复制
发表时间:
2008-09-01
期刊:
影响因子:
5.8
通讯作者:
Craig, David
Craig, David
中科院分区:
生物学3区
文献类型:
--
作者:
Homer, Nils;Tembe, Waibhav D.;Craig, David

文献摘要

被引文献

相似文献

对于许多全基因组关联(GWA)研究,单独地对一百万或更多的SNP进行基因分型以大量成本提供覆盖率的边际增加。由于人类基因组固有的相关结构,获得的许多信息都是冗余的。基于池化的GWA研究可以通过利用这种冗余来减少噪音,提高观察的准确性并增加基因组覆盖率而显着受益。我们引入了个体基因分型和合并之间相关性的度量,在相同的框架下,r(2)提供了SNP对之间连锁不平衡(LD)的度量。然后,我们报告了一种新的非单倍型多标记多位点方法,该方法利用人类基因组中SNP之间的相关结构来提高基于池的GWA研究的效率。我们首先给出了我们的多标记方法的理论框架和推导。接下来,我们评估模拟使用这种多标记的方法相比,单标记分析。最后,我们在Illumina 450S Duo、Illumina 550K和Affyssin 5.0平台上使用不同的HapMap个体池对我们的方法进行了实验评估,总共有1333631个SNP。我们的研究结果表明,使用多标记分析减少了特定于基于池的研究的噪声,允许多个微阵列平台的有效整合,并提供比单标记分析更准确的显著性测量。此外,这种方法可以扩展到允许使用LD中的相邻SNP直接观察到的SNP的关联显著性。这种多标记方法现在可以用于在超过一百万个SNP的多个平台上经济高效地完成基于池的GWA研究,并对由于池化而导致的信息损失进行加权的相邻SNP进行估算。
For many genome-wide association (GWA) studies individually genotyping one million or more SNPs provides a marginal increase in coverage at a substantial cost. Much of the information gained is redundant due to the correlation structure inherent in the human genome. Pooling-based GWA studies could benefit significantly by utilizing this redundancy to reduce noise, improve the accuracy of the observations and increase genomic coverage. We introduce a measure of correlation between individual genotyping and pooling, under the same framework that r(2) provides a measure of linkage disequilibrium (LD) between pairs of SNPs. We then report a new non-haplotype multimarker multi-loci method that leverages the correlation structure between SNPs in the human genome to increase the efficacy of pooling-based GWA studies. We first give a theoretical framework and derivation of our multimarker method. Next, we evaluate simulations using this multimarker approach in comparison to single marker analysis. Finally, we experimentally evaluate our method using different pools of HapMap individuals on the Illumina 450S Duo, Illumina 550K and Affymetrix 5.0 platforms for a combined total of 1 333 631 SNPs. Our results show that use of multimarker analysis reduces noise specific to pooling-based studies, allows for efficient integration of multiple microarray platforms and provides more accurate measures of significance than single marker analysis. Additionally, this approach can be extended to allow for imputing the association significance for SNPs not directly observed using neighboring SNPs in LD. This multimarker method can now be used to cost-effectively complete pooling-based GWA studies with multiple platforms across over one million SNPs and to impute neighboring SNPs weighted for the loss of information due to pooling.