Corrigendum of 'High throughput analysis of epistasis in genome-wide association studies with BiForce'

Corrigendum of 'High throughput analysis of epistasis in genome-wide association studies with BiForce'
复制标题

“BiForce 全基因组关联研究中上位性的高通量分析”勘误表

DOI:
10.1093/bioinformatics/btt444
复制
发表时间:
2013
期刊:
影响因子:
5.8
通讯作者:
Gyenesei A
Gyenesei A
中科院分区:
生物学3区
文献类型:
--
作者:
Gyenesei A

文献摘要

相似文献

在我们的文章发表之后,描述了BiForce的使用(http://bioinfo. utu。fi/biforcetoolbox)用于上位性分析(Gyenesei等人,2012),我们观察到,在分析数量性状时,在特殊情况下可能会出现上位性的夸大证据。这可能发生在一个上位性对中两个相邻的(例如相距<200kb)单核苷酸多态性(SNP)处于连锁不平衡(LD),并且其中至少一个携带强的边缘效应。最近在其他基于LD或单倍型的方法中发现了类似的膨胀,用于分析疾病性状中的上位性(Ueki和Cordell,2012)。在我们的情况下,这个问题不影响疾病性状的分析,因为BiForce使用逻辑回归作为最后一步来生成这些性状的结果(Wan等人,2010年)。因此,原始论文的表3(Gyenesei等人,2012)是正确的。然而,对于数量性状,BiForce使用基于列联表的F比率检验进行相互作用,而无需线性回归中应用的拟合步骤。众所周知,这些测试不是正交的,但当两个SNP之间的LD较低时,它们是稳健的,允许通过BiForce实现快速筛选。然而,当LD高时,两个相关SNP之间的相互作用的测试被SNP对的边际效应膨胀,因此当边际效应强时膨胀是关键的,但当边际效应弱时不是。然而,由于BiForce默认使用严格的Bonferroni调整阈值,因此一般来说,膨胀上位对在全基因组范围内具有显著性的机会应该很低。该问题影响了原始文章表2中的结果(Gyenesei等人,2012年)。每个SNP对的正确相互作用P值(Pint)列于本文后面更新的表2中,表明在C-反应蛋白(CRP)、葡萄糖(GLU)、高密度脂蛋白(HDL)、低密度脂蛋白(LDL)和甘油三酯(TRI)中没有一个保持全基因组显著性。对数量性状和疾病性状的模拟数据的分析结果不受影响,因为在模拟中,SNPs是从染色体中随机抽取的,假设它们处于Hardy-Weinberg平衡,即高LD与强边缘效应一起出现的机会非常低,这得到了原始文章图3中假阳性率结果的支持。总而言之,该问题只影响一小部分
Following the publication of our article, describing the use of BiForce (http://bioinfo. utu. fi/biforcetoolbox) for the analysis of epistasis (Gyenesei et al., 2012), we observed that inflated evidence for epistasis may arise under exceptional circumstances when analyzing quantitative traits. This may occur when two neighboring (eg< 200kb apart) single nucleotide polymorphisms (SNPs) in an epistatic pair are in linkage disequilibrium (LD) and at least one of them carries strong marginal effects. Similar inflation was discovered recently in other LD-or haplotype-based methods for the analysis of epistasis in disease traits (Ueki and Cordell, 2012). This issue does not affect the analysis of disease traits in our case because BiForce uses logistic regression as the final step to generate the results for such traits (Wan et al., 2010). Thus, Table 3 of the original paper (Gyenesei et al., 2012) is correct. However for quantitative traits, BiForce uses contingency table-based F ratio tests for interactions without the fitting step applied in linear regression. It is known that such tests are not orthogonal, but they are robust when LD between two SNPs is low, allowing the fast screening achieved by BiForce. When LD is high, however, the test for interaction between two correlated SNPs is inflated by the marginal effects of the pair of SNPs, and therefore the inflation is critical when marginal effects are strong but not when marginal effects are weak. Nevertheless, because BiForce uses stringent Bonferroni-adjusted thresholds by default, the chance of inflated epistatic pairs being genomewide significant should be low in general. This issue affected the results in Table 2 of the original article (Gyenesei et al., 2012). The correct interaction P-values (Pint) of each SNP pair are listed in the updated Table 2 later in the text, suggesting that none remained genome-wide significant in C-reactive protein (CRP), glucose (GLU), high-density lipoprotein (HDL), low-density lipoprotein (LDL) and triglycerides (TRI). Results from the analyses of simulated data on quantitative and disease traits are unaffected because in simulation SNPs were randomly drawn from a chromosome assuming they were in Hardy–Weinberg equilibrium, ie the chance of high LD coming together with strong marginal effects is very low, which is supported by the results of false positive rate in Figure 3 of the original article. In summary, the issue only affects a small part of