SNP imputation bias reduces effect size determination.

SNP imputation bias reduces effect size determination.
复制标题

DOI:
10.3389/fgene.2015.00030
复制
发表时间:
2015
影响因子:
3.7
通讯作者:
Baranzini SE
Baranzini SE
中科院分区:
生物学3区
文献类型:
--
作者:
Khankhanian P;Din L;Caillier SJ;Gourraud PA;Baranzini SE

文献摘要

被引文献

相似文献

插补是一种常用的技术,它利用连锁不平衡,使用特征明确的参考群体来推断遗传数据集中缺失的基因型。虽然人们一致认为参考人群必须与查询数据集的种族相匹配,但通常的做法是使用相同的参考来估算各种表型的基因型。我们假设使用由具有与查询数据集不同表型的样本组成的参考会引入插补偏差。为了检验这一假设,我们使用了肌萎缩侧索硬化症 (ALS)、帕金森病 (PD) 和克罗恩病 (CD) 的 GWAS 数据集。首先,我们对每项研究中的 100 个疾病相关标记和 100 个非相关标记进行掩蔽和插补。并行使用两个估算参考:一个由健康对照组成,另一个由患有相同疾病的患者组成。我们通过将预测的基因型与 SNP 芯片检测的基因型进行比较来评估插补的不一致(不精确性)和偏差(不准确)。我们还评估了在 GWAS 研究中使用预测基因型时观察到的效应大小的偏差。当使用健康对照作为插补的参考时,观察到显着的偏差,特别是在与疾病相关的标记物中。使用案例作为参考大大减弱了这种偏见。对于几乎所有标记,偏差的方向有利于非风险等位基因。在这三种疾病的 GWAS 研究中(以 1000 个基因组的健康参考对照作为参考),通过插补获得的疾病相关标记的平均 OR 低于使用原始分析基因型获得的平均 OR。我们发现偏差是插补所固有的,因为使用不同的方法不会改变结果。总之,插补是预测 GWAS 基因型和估计遗传风险的有效方法。然而,需要仔细选择参考人群,以尽量减少这种方法固有的偏差。
Imputation is a commonly used technique that exploits linkage disequilibrium to infer missing genotypes in genetic datasets, using a well-characterized reference population. While there is agreement that the reference population has to match the ethnicity of the query dataset, it is common practice to use the same reference to impute genotypes for a wide variety of phenotypes. We hypothesized that using a reference composed of samples with a different phenotype than the query dataset would introduce imputation bias. To test this hypothesis we used GWAS datasets from Amyotrophic Lateral Sclerosis (ALS), Parkinson Disease (PD), and Crohn's Disease (CD). First, we masked and then performed imputation of 100 disease-associated markers and 100 non-associated markers from each study. Two references for imputation were used in parallel: one consisting of healthy controls and another consisting of patients with the same disease. We assessed the discordance (imprecision) and bias (inaccuracy) of imputation by comparing predicted genotypes to those assayed by SNP-chip. We also assessed the bias on the observed effect size when the predicted genotypes were used in a GWAS study. When healthy controls were used as reference for imputation, a significant bias was observed, particularly in the disease-associated markers. Using cases as reference significantly attenuated this bias. For nearly all markers, the direction of the bias favored the non-risk allele. In GWAS studies of the three diseases (with healthy reference controls from the 1000 genomes as reference), the mean OR for disease-associated markers obtained by imputation was lower than that obtained using original assayed genotypes. We found that the bias is inherent to imputation as using different methods did not alter the results. In conclusion, imputation is a powerful method to predict genotypes and estimate genetic risk for GWAS. However, a careful choice of reference population is needed to minimize biases inherent to this approach.