A benchmark study on current GWAS models in admixed populations.

A benchmark study on current GWAS models in admixed populations.
复制标题

混合人群中当前GWAS模型的基准研究。

DOI:
10.1093/bib/bbad437
复制
发表时间:
2023-11-22
影响因子:
9.5
通讯作者:
--
中科院分区:
生物学2区
文献类型:
--
作者:

文献摘要

参考文献

相似文献

流行的全基因组关联研究 (GWAS) 模型的性能尚未在遗传混合场景下以一致的方式进行检验,这引入了几个具有挑战性的方面:次要等位基因频率 (MAF) 的异质性、广泛的病例对照比、不同的效应大小等。我们生成了一组合成个体 (N = 19 234),用于模拟 (i) 大样本量; (ii) 双向混合(美洲原住民和欧洲血统)和 (iii) 二元表型。然后,我们通过计算不同 MAF、病例对照比率、样本大小和不同血统比例下的通货膨胀因子和功效计算,对三种流行的 GWAS 工具 [广义线性混合模型相关测试 (GMMAT)、广义混合模型 (SAIGE) 的可扩展且准确的实现和 Tractor] 进行基准测试。我们还雇佣了一群秘鲁人 (N = 249) 来进一步检查测试模型在 (i) 真实遗传和表型数据以及 (ii) 小样本量上的表现。在综合队列中,SAIGE 在 I 类错误率方面表现优于 GMMAT 和 Tractor,特别是在病例对照比严重失衡的情况下。相反,功效分析认为 Tractor 是查明祖先特定因果变异的最佳方法,但当效应大小显示祖先之间的异质性有限时,功效就会下降。在秘鲁队列中,只有 Tractor 识别出两个与美洲原住民血统相关的暗示性位点(P 值)。当前的研究说明了遗传混合情况下可用 GWAS 工具的最佳实践和局限性。尽管需要仔细考虑复杂的情况(小样本量、不平衡的病例对照比、MAF 异质性),但将当地血统纳入 GWAS 分析可以增强功效。
The performances of popular genome-wide association study (GWAS) models have not been examined yet in a consistent manner under the scenario of genetic admixture, which introduces several challenging aspects: heterogeneity of minor allele frequency (MAF), wide spectrum of case–control ratio, varying effect sizes, etc. We generated a cohort of synthetic individuals (N = 19 234) that simulates (i) a large sample size; (ii) two-way admixture (Native American and European ancestry) and (iii) a binary phenotype. We then benchmarked three popular GWAS tools [generalized linear mixed model associated test (GMMAT), scalable and accurate implementation of generalized mixed model (SAIGE) and Tractor] by computing inflation factors and power calculations under different MAFs, case–control ratios, sample sizes and varying ancestry proportions. We also employed a cohort of Peruvians (N = 249) to further examine the performances of the testing models on (i) real genetic and phenotype data and (ii) small sample sizes. In the synthetic cohort, SAIGE performed better than GMMAT and Tractor in terms of type-I error rate, especially under severe unbalanced case–control ratio. On the contrary, power analysis identified Tractor as the best method to pinpoint ancestry-specific causal variants but showed decreased power when the effect size displayed limited heterogeneity between ancestries. In the Peruvian cohort, only Tractor identified two suggestive loci (P-value ) associated with Native American ancestry. The current study illustrates best practice and limitations for available GWAS tools under the scenario of genetic admixture. Incorporating local ancestry in GWAS analyses boosts power, although careful consideration of complex scenarios (small sample sizes, imbalance case–control ratio, MAF heterogeneity) is needed.
DOI: 10.1038/s41588-020-00766-y
发表时间: 2021-03
期刊: Nature genetics
影响因子: 30.8
作者:
Atkinson EG;Maihofer AX;Kanai M;Martin AR;Karczewski KJ;Santoro ML;Ulirsch JC;Kamatani Y;Okada Y;Finucane HK;Koenen KC;Nievergelt CM;Daly MJ;Neale BM
通讯作者: Neale BM
DOI: 10.1186/s13742-015-0047-8
发表时间: 2015
期刊: GigaScience
影响因子: 9.2
作者:
Chang CC;Chow CC;Tellier LC;Vattikuti S;Purcell SM;Lee JJ
通讯作者: Lee JJ
DOI: 10.1093/biomet/86.4.929
发表时间: 1999-12-01
期刊: BIOMETRIKA
影响因子: 2.7
作者:
Kuonen, D
通讯作者: Kuonen, D
DOI: 10.1002/gepi.22359
发表时间: 2020-10-22
影响因子: 2.1
作者:
Sofer, Tamar;Guo, Na
通讯作者: Guo, Na
DOI: 10.1016/j.ajhg.2013.06.020
发表时间: 2013-08-08
影响因子: 9.8
作者:
Maples, Brian K.;Gravel, Simon;Bustamante, Carlos D.
通讯作者: Bustamante, Carlos D.