Optimal population-specific HLA imputation with dimension reduction

Optimal population-specific HLA imputation with dimension reduction
复制标题

DOI:
10.1111/tan.15282
复制
发表时间:
2023-11-11
期刊:
HLA
影响因子:
8
通讯作者:
Vince,Nicolas
Vince,Nicolas
中科院分区:
医学4区
文献类型:
--
作者:
Douillard,Venceslas;Silva,Nayane dos Santos Brito;Vince,Nicolas

文献摘要

被引文献

相似文献

人类基因组学发展迅速,为全基因组关联研究(GWAS)提供了动力。基于SNP的GWAS不能捕捉与疾病易感性高度相关的HLA基因的强烈多态性。有一些方法可以从SNP基因数据中统计出HLA基因类型,但缺乏参考面板的多样性阻碍了它们的表现。我们评估了1000个基因组数据作为参考面板的准确性,以从非洲和欧洲祖先的混合个体中输入人类白细胞抗原,重点是(A)完整的数据集,(B)来自6个群体的10个重复,以及(C)定制参考面板的19个条件。完整的数据集表现优于较小的模型,对于HLA-B,F1得分为0.66。然而,定制模型的表现优于类似大小的多种族或人群模型(F1得分高达0.53,而不是高达0.42)。我们证明了使用遗传特异性模型来推算目前在公共数据集中未被充分代表的群体的重要性,从而为每个遗传群体的人类白细胞抗原推算打开了大门。
Human genomics has quickly evolved, powering genome‐wide association studies (GWASs). SNP‐based GWASs cannot capture the intense polymorphism ofHLAgenes, highly associated with disease susceptibility. There are methods to statistically imputeHLAgenotypes from SNP‐genotypes data, but lack of diversity in reference panels hinders their performance. We evaluated the accuracy of the 1000 Genomes data as a reference panel for imputing HLA from admixed individuals of African and European ancestries, focusing on (a) the full dataset, (b) 10 replications from 6 populations, and (c) 19 conditions for the custom reference panels. The full dataset outperformed smaller models, with a good F1‐score of 0.66 forHLA‐B. However, custom models outperformed the multiethnic or population models of similar size (F1‐scores up to 0.53, against up to 0.42). We demonstrated the importance of using genetically specific models for imputing populations, which are currently underrepresented in public datasets, opening the door to HLA imputation for every genetic population.