Genotype imputation with thousands of genomes.

Genotype imputation with thousands of genomes.
复制标题

DOI:
10.1534/g3.111.001198
复制
发表时间:
2011-11
期刊:
G3 (Bethesda, Md.)
影响因子:
--
通讯作者:
Stephens M
Stephens M
中科院分区:
其他
文献类型:
--
作者:
Howie B;Marchini J;Stephens M

文献摘要

参考文献

被引文献

相似文献

基因型插补是一种统计技术,常用于提高遗传关联研究的功效和分辨率。插补方法通过使用参考组中的单倍型模式来预测研究数据集中未观察到的基因型,并且已经提出了许多方法来选择参考单倍型的子集,以最大限度地提高给定研究人群的准确性。这些小组选择策略变得更难应用和解释,因为像1000个基因组计划这样的测序工作产生了更大和更多样化的参考集,这促使我们开发了一种替代框架。我们的方法是建立在一个新的近似,使用本地序列相似性选择一个自定义的参考面板为每个研究单倍型在基因组的每个区域。这种近似使得使用所有可用的参考单倍型在计算上是有效的,这使得我们能够绕过小组选择步骤,并通过捕获群体之间意外的等位基因共享来提高低频变异的准确性。使用HapMap 3的数据,我们表明我们的框架在广泛的人群中产生了准确的结果。我们还使用疟疾遗传流行病学网络(MalariaGEN)的数据为非洲基于估算的研究提供建议。我们证明了我们的近似提高了大型基于序列的参考面板的效率,并且我们讨论了现代参考数据集的通用计算策略。全基因组关联研究将很快能够利用数千个参考基因组的力量,我们的工作为研究人员使用这些丰富的信息提供了一种实用的方法。在IMPUTE2软件包中实施了本研究的新方法。
Genotype imputation is a statistical technique that is often used to increase the power and resolution of genetic association studies. Imputation methods work by using haplotype patterns in a reference panel to predict unobserved genotypes in a study dataset, and a number of approaches have been proposed for choosing subsets of reference haplotypes that will maximize accuracy in a given study population. These panel selection strategies become harder to apply and interpret as sequencing efforts like the 1000 Genomes Project produce larger and more diverse reference sets, which led us to develop an alternative framework. Our approach is built around a new approximation that uses local sequence similarity to choose a custom reference panel for each study haplotype in each region of the genome. This approximation makes it computationally efficient to use all available reference haplotypes, which allows us to bypass the panel selection step and to improve accuracy at low-frequency variants by capturing unexpected allele sharing among populations. Using data from HapMap 3, we show that our framework produces accurate results in a wide range of human populations. We also use data from the Malaria Genetic Epidemiology Network (MalariaGEN) to provide recommendations for imputation-based studies in Africa. We demonstrate that our approximation improves efficiency in large, sequence-based reference panels, and we discuss general computational strategies for modern reference datasets. Genome-wide association studies will soon be able to harness the power of thousands of reference genomes, and our work provides a practical way for investigators to use this rich information. New methodology from this study is implemented in the IMPUTE2 software package.
DOI: 10.1371/journal.pgen.1000279
发表时间: 2008-12
期刊: PLOS GENETICS
影响因子: 4.5
作者:
Guan, Yongtao;Stephens, Matthew
通讯作者: Stephens, Matthew
DOI: 10.1093/bioinformatics/btn522
发表时间: 2008-12-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Hernandez, Ryan D.
通讯作者: Hernandez, Ryan D.
DOI: 10.1016/j.cub.2009.11.050
发表时间: 2010-02-23
期刊: Current biology : CB
影响因子: --
作者:
Campbell MC;Tishkoff SA
通讯作者: Tishkoff SA
DOI: 10.1038/ng2088
发表时间: 2007-07-01
期刊: NATURE GENETICS
影响因子: 30.8
作者:
Marchini, Jonathan;Howie, Bryan;Donnelly, Peter
通讯作者: Donnelly, Peter
DOI: 10.1038/ejhg.2011.10
发表时间: 2011-06-01
影响因子: 5.2
作者:
Jostins, Luke;Morley, Katherine I.;Barrett, Jeffrey C.
通讯作者: Barrett, Jeffrey C.