PRIMAL: Fast and accurate pedigree-based imputation from sequence data in a founder population.
PRIMAL: Fast and accurate pedigree-based imputation from sequence data in a founder population.
复制标题
DOI:
10.1371/journal.pcbi.1004139
复制
发表时间:
2015-03
影响因子:
4.3
通讯作者:
Nicolae DL
中科院分区:
文献类型:
--
作者:
Livne OE;Han L;Alkorta-Aranburu G;Wentworth-Sheilds W;Abney M;Ober C;Nicolae DL
Founder populations and large pedigrees offer many well-known advantages for genetic mapping studies, including cost-efficient study designs. Here, we describe PRIMAL (PedigRee IMputation ALgorithm), a fast and accurate pedigree-based phasing and imputation algorithm for founder populations. PRIMAL incorporates both existing and original ideas, such as a novel indexing strategy of Identity-By-Descent (IBD) segments based on clique graphs. We were able to impute the genomes of 1,317 South Dakota Hutterites, who had genome-wide genotypes for ~300,000 common single nucleotide variants (SNVs), from 98 whole genome sequences. Using a combination of pedigree-based and LD-based imputation, we were able to assign 87% of genotypes with >99% accuracy over the full range of allele frequencies. Using the IBD cliques we were also able to infer the parental origin of 83% of alleles, and genotypes of deceased recent ancestors for whom no genotype information was available. This imputed data set will enable us to better study the relative contribution of rare and common variants on human phenotypes, as well as parental origin effect of disease risk alleles in >1,000 individuals at minimal cost. The recent availability of whole genome and whole exome sequencing allows genetic studies of human diseases and traits at an unprecedented resolution, although their cost limits the size of the studied sample. To overcome this limitation and design cost-efficient studies, we developed a two step method: sequencing of relatively few members of a well-characterized founder population followed by pedigree-based whole genome imputation of many other individuals with genome-wide genotype data. We show that by sequencing only 98 Hutterites, we can impute 7 million variants in an additional 1,317 Hutterites with >99% accuracy and an average call rate of 87%. Furthermore, parental origin was assigned to 83% of the alleles. Such studies in the Hutterites and other founder populations should yield new insights into the genetic architecture of common diseases, gene expression traits, and clinically relevant biomarkers of disease, and ultimately provide outstanding opportunities for personalized medicine in these well-characterized populations.
登录
查看更多内容
影响因子:
3.7
作者:
Li L;Li Y;Browning SR;Browning BL;Slater AJ;Kong X;Aponte JL;Mooser VE;Chissoe SL;Whittaker JC;Nelson MR;Ehm MG
通讯作者:
Ehm MG
影响因子:
9.8
作者:
Cheung, Charles Y. K.;Thompson, Elizabeth A.;Wijsman, Ellen M.
通讯作者:
Wijsman, Ellen M.
影响因子:
3.1
作者:
Livne, Oren E.;Brandt, Achi
通讯作者:
Brandt, Achi
影响因子:
2.1
作者:
Li, Yun;Willer, Cristen J.;Ding, Jun;Scheet, Paul;Abecasis, Goncalo R.
通讯作者:
Abecasis, Goncalo R.
影响因子:
30.8
作者:
Abecasis, GR;Cherny, SS;Cardon, LR
通讯作者:
Cardon, LR