Improving genetic risk prediction across diverse population by disentangling ancestry representations.

Improving genetic risk prediction across diverse population by disentangling ancestry representations.
复制标题

DOI:
10.1038/s42003-023-05352-6
复制
发表时间:
2023-09-22
影响因子:
5.9
通讯作者:
He, Zihuai
He, Zihuai
中科院分区:
生物学2区
文献类型:
--
作者:
Gyawali, Prashnna K.;Le Guen, Yann;Liu, Xiaoxia;Belloy, Michael E.;Tang, Hua;Zou, James;He, Zihuai

文献摘要

参考文献

相似文献

使用遗传数据的风险预测模型在基因组学中的吸引力越来越大。然而,大多数多基因风险模型是使用具有相似血统(主要是欧洲人)的参与者的数据开发的。这可能会导致风险预测因子的偏差,导致在应用于少数族裔人口和非洲裔美国人等混杂个人时通用性较差。为了解决这个问题,很大程度上是由于预测模型受到潜在人口结构的偏差,我们提出了一个深度学习框架,该框架利用来自不同人群的数据,并在其表示中从与表型相关的信息中分离祖先。血统分离的表示法可以用来建立风险预测因子,在少数族裔人群中表现得更好。我们将所提出的方法应用于阿尔茨海默病的遗传学分析。与标准的线性和非线性风险预测方法相比,该方法显著改善了少数民族人群(包括混血个体)的风险预测,而不需要自我报告祖先信息。深度学习框架利用来自不同人群的数据,并在其表示中将祖先与与表型相关的信息分开。
Risk prediction models using genetic data have seen increasing traction in genomics. However, most of the polygenic risk models were developed using data from participants with similar (mostly European) ancestry. This can lead to biases in the risk predictors resulting in poor generalization when applied to minority populations and admixed individuals such as African Americans. To address this issue, largely due to the prediction models being biased by the underlying population structure, we propose a deep-learning framework that leverages data from diverse population and disentangles ancestry from the phenotype-relevant information in its representation. The ancestry disentangled representation can be used to build risk predictors that perform better across minority populations. We applied the proposed method to the analysis of Alzheimer’s disease genetics. Comparing with standard linear and nonlinear risk prediction methods, the proposed method substantially improves risk prediction in minority populations, including admixed individuals, without needing self-reported ancestry information. A deep-learning framework leverages data from diverse populations and disentangles ancestry from the phenotype-relevant information in its representation.
DOI: 10.1212/nxg.0000000000000141
发表时间: 2017-04
期刊: Neurology. Genetics
影响因子: --
作者:
N'Songo A;Carrasquillo MM;Wang X;Burgess JD;Nguyen T;Asmann YW;Serie DJ;Younkin SG;Allen M;Pedraza O;Duara R;Greig Custo MT;Graff-Radford NR;Ertekin-Taner N
通讯作者: Ertekin-Taner N
DOI: 10.1038/s41467-023-36544-7
发表时间: 2023-02-14
影响因子: 16.6
作者:
Miao, Jiacheng;Guo, Hanmin;Song, Gefei;Zhao, Zijie;Hou, Lin;Lu, Qiongshi
通讯作者: Lu, Qiongshi
DOI: 10.1038/s41588-020-00766-y
发表时间: 2021-03
期刊: Nature genetics
影响因子: 30.8
作者:
Atkinson EG;Maihofer AX;Kanai M;Martin AR;Karczewski KJ;Santoro ML;Ulirsch JC;Kamatani Y;Okada Y;Finucane HK;Koenen KC;Nievergelt CM;Daly MJ;Neale BM
通讯作者: Neale BM
DOI: 10.1016/j.neurobiolaging.2016.07.018
发表时间: 2017-01-01
影响因子: 4.2
作者:
Escott-Price, Valentina;Shoai, Maryam;Hardy, John
通讯作者: Hardy, John
DOI: 10.1038/ejhg.2016.17
发表时间: 2016-08
期刊: European journal of human genetics : EJHG
影响因子: --
作者:
Cook JP;Morris AP
通讯作者: Morris AP