Deep learning-based phenotype imputation on population-scale biobank data increases genetic discoveries.

Deep learning-based phenotype imputation on population-scale biobank data increases genetic discoveries.
复制标题

DOI:
10.1038/s41588-023-01558-w
复制
发表时间:
2023-12
期刊:
影响因子:
30.8
通讯作者:
Sankararaman, Sriram
Sankararaman, Sriram
中科院分区:
生物学1区
文献类型:
--
作者:
An, Ulzee;Pazokitoroudi, Ali;Alvarez, Marcus;Huang, Lianyun;Bacanu, Silviu;Schork, Andrew J.;Kendler, Kenneth;Pajukanta, Paeivi;Flint, Jonathan;Zaitlen, Noah;Cai, Na;Dahl, Andy;Sankararaman, Sriram

文献摘要

参考文献

被引文献

相似文献

收集许多个体的深层表型和基因组数据的生物库已成为人类遗传学的关键资源。然而,生物库中的表型往往在许多个体中缺失,限制了它们的实用性。我们提出了AutoComplete,这是一种基于深度学习的插补方法,用于在人口规模的生物库数据集中插补或“填充”缺失的表型。当应用于从英国生物库测量的约30万个体的表型集合时,AutoComplete大大提高了现有方法的插补准确性。在三个具有显著缺失量的性状上,我们表明AutoComplete产生与最初观察到的表型在遗传上相似的估算表型,同时将有效样本量平均增加约两倍。此外,对所得插补表型的全基因组关联分析导致相关基因座数量大幅增加。我们的研究结果证明了基于深度学习的表型插补在现有生物库数据集中提高遗传发现能力的实用性。AutoComplete是一种基于深度学习的方法,可以在人群规模的生物库数据集中估算缺失的表型,增加有效样本量,并提高全基因组关联研究中遗传发现的能力。
Biobanks that collect deep phenotypic and genomic data across many individuals have emerged as a key resource in human genetics. However, phenotypes in biobanks are often missing across many individuals, limiting their utility. We propose AutoComplete, a deep learning-based imputation method to impute or ‘fill-in’ missing phenotypes in population-scale biobank datasets. When applied to collections of phenotypes measured across ~300,000 individuals from the UK Biobank, AutoComplete substantially improved imputation accuracy over existing methods. On three traits with notable amounts of missingness, we show that AutoComplete yields imputed phenotypes that are genetically similar to the originally observed phenotypes while increasing the effective sample size by about twofold on average. Further, genome-wide association analyses on the resulting imputed phenotypes led to a substantial increase in the number of associated loci. Our results demonstrate the utility of deep learning-based phenotype imputation to increase power for genetic discoveries in existing biobank datasets. AutoComplete is a deep learning-based method that imputes missing phenotypes in population-scale biobank datasets, increasing effective sample sizes and improving power for genetic discoveries in genome-wide association studies.
DOI: 10.1038/ng2088
发表时间: 2007-07-01
期刊: NATURE GENETICS
影响因子: 30.8
作者:
Marchini, Jonathan;Howie, Bryan;Donnelly, Peter
通讯作者: Donnelly, Peter
DOI: 10.1186/s13073-020-00820-8
发表时间: 2021-01-13
期刊: Genome medicine
影响因子: 12.3
作者:
Dennis JK;Sealock JM;Straub P;Lee YH;Hucks D;Actkins K;Faucon A;Feng YA;Ge T;Goleva SB;Niarchou M;Singh K;Morley T;Smoller JW;Ruderfer DM;Mosley JD;Chen G;Davis LK
通讯作者: Davis LK
DOI: 10.1038/s41586-018-0579-z
发表时间: 2018-10
期刊: Nature
影响因子: 64.8
作者:
Bycroft C;Freeman C;Petkova D;Band G;Elliott LT;Sharp K;Motyer A;Vukcevic D;Delaneau O;O'Connell J;Cortes A;Welsh S;Young A;Effingham M;McVean G;Leslie S;Allen N;Donnelly P;Marchini J
通讯作者: Marchini J
DOI: 10.1038/ng.3513
发表时间: 2016-04
期刊: Nature genetics
影响因子: 30.8
作者:
Dahl A;Iotchkova V;Baud A;Johansson Å;Gyllensten U;Soranzo N;Mott R;Kranis A;Marchini J
通讯作者: Marchini J
DOI: 10.1038/ng.3211
发表时间: 2015-03
期刊: NATURE GENETICS
影响因子: 30.8
作者:
Bulik-Sullivan, Brendan K.;Loh, Po-Ru;Finucane, Hilary K.;Ripke, Stephan;Yang, Jian;Patterson, Nick;Daly, Mark J.;Price, Alkes L.;Neale, Benjamin M.
通讯作者: Neale, Benjamin M.