The hazards of genotype imputation when mapping disease susceptibility variants.

The hazards of genotype imputation when mapping disease susceptibility variants.
复制标题

DOI:
10.1186/s13059-023-03140-3
复制
发表时间:
2024-01-03
期刊:
影响因子:
12.3
通讯作者:
--
中科院分区:
生物学1区
文献类型:
--
作者:

文献摘要

参考文献

相似文献

使用插补来推断缺失基因型的统计能力的无成本增加无疑很有吸引力,但它是否没有风险?这个针对 3 个 2 型糖尿病 (T2D) 基因座的案例研究表明事实并非如此;它揭示了为什么会出现这种情况,并引起了人们对疾病位点估算缺陷的担忧,其中病例和参考组之间的单倍型不同。之前使用靶向测序鉴定了 T2D 相关变异。我们删除了这些显着相关的 SNP,并使用邻近的 SNP 通过插补来推断它们。我们将估算的基因型与观察到的基因型进行比较,检查 T2D-SNP 关联的改变模式,并通过研究单倍型结构来调查估算错误的原因。大多数 T2D 变异都被错误地用低密度的支架 SNP 进行了估算,但大多数即使在高密度下也未能进行估算,尽管获得了高确定性分数。对于风险等位基因,观察到的缺失和不一致的插补错误不成比例,产生单态基因型识别或假阴性关联。我们表明,对于所有位点,携带风险等位基因的单倍型在 T2D 病例中比参考组更为常见。插补并不是精细绘图的万能药,也不是基于不同阵列和不同群体对多个 GWAS 进行荟萃分析的万能药。我们测试的总共 80% 的 SNP 不包含在阵列平台中,这解释了为什么这些和其他此类相关变体以前可能被遗漏。无论软件和参考单倍型的选择如何,插补都会推动对参考组的基因型推断,从而在疾病位点引入错误。在线版本包含可在 10.1186/s13059-023-03140-3 获取的补充材料。
The cost-free increase in statistical power of using imputation to infer missing genotypes is undoubtedly appealing, but is it hazard-free? This case study of three type-2 diabetes (T2D) loci demonstrates that it is not; it sheds light on why this is so and raises concerns as to the shortcomings of imputation at disease loci, where haplotypes differ between cases and reference panel. T2D-associated variants were previously identified using targeted sequencing. We removed these significantly associated SNPs and used neighbouring SNPs to infer them by imputation. We compared imputed with observed genotypes, examined the altered pattern of T2D-SNP association, and investigated the cause of imputation errors by studying haplotype structure. Most T2D variants were incorrectly imputed with a low density of scaffold SNPs, but the majority failed to impute even at high density, despite obtaining high certainty scores. Missing and discordant imputation errors, which were observed disproportionately for the risk alleles, produced monomorphic genotype calls or false-negative associations. We show that haplotypes carrying risk alleles are considerably more common in the T2D cases than the reference panel, for all loci. Imputation is not a panacea for fine mapping, nor for meta-analysing multiple GWAS based on different arrays and different populations. A total of 80% of the SNPs we have tested are not included in array platforms, explaining why these and other such associated variants may previously have been missed. Regardless of the choice of software and reference haplotypes, imputation drives genotype inference towards the reference panel, introducing errors at disease loci. The online version contains supplementary material available at 10.1186/s13059-023-03140-3.
DOI: 10.1159/000489758
发表时间: 2017-01-01
期刊: HUMAN HEREDITY
影响因子: 1.8
作者:
Shi, Shuo;Yuan, Na;Xiao, Jingfa
通讯作者: Xiao, Jingfa
DOI: 10.1038/ng.871
发表时间: 2011-07-24
期刊: Nature genetics
影响因子: 30.8
作者:
通讯作者: --
DOI: 10.3389/fgene.2015.00030
发表时间: 2015
影响因子: 3.7
作者:
Khankhanian P;Din L;Caillier SJ;Gourraud PA;Baranzini SE
通讯作者: Baranzini SE
DOI: 10.1002/gepi.20185
发表时间: 2007-11-01
影响因子: 2.1
作者:
Andres, Aida M.;Clark, Andrew G.;Hixson, James E.
通讯作者: Hixson, James E.
DOI: 10.1007/s00125-006-0502-2
发表时间: 2007-01-01
期刊: DIABETOLOGIA
影响因子: 8.2
作者:
Chandak, G. R.;Janipalli, C. S.;Yajnik, C. S.
通讯作者: Yajnik, C. S.