Corrigendum of 'High throughput analysis of epistasis in genome-wide association studies with BiForce'
Corrigendum of 'High throughput analysis of epistasis in genome-wide association studies with BiForce'
复制标题
“BiForce 全基因组关联研究中上位性的高通量分析”勘误表
DOI:
10.1093/bioinformatics/btt444
复制
发表时间:
2013
期刊:
影响因子:
5.8
通讯作者:
Gyenesei A
中科院分区:
文献类型:
--
作者:
Gyenesei A
Following the publication of our article, describing the use of BiForce (http://bioinfo. utu. fi/biforcetoolbox) for the analysis of epistasis (Gyenesei et al., 2012), we observed that inflated evidence for epistasis may arise under exceptional circumstances when analyzing quantitative traits. This may occur when two neighboring (eg< 200kb apart) single nucleotide polymorphisms (SNPs) in an epistatic pair are in linkage disequilibrium (LD) and at least one of them carries strong marginal effects. Similar inflation was discovered recently in other LD-or haplotype-based methods for the analysis of epistasis in disease traits (Ueki and Cordell, 2012). This issue does not affect the analysis of disease traits in our case because BiForce uses logistic regression as the final step to generate the results for such traits (Wan et al., 2010). Thus, Table 3 of the original paper (Gyenesei et al., 2012) is correct. However for quantitative traits, BiForce uses contingency table-based F ratio tests for interactions without the fitting step applied in linear regression. It is known that such tests are not orthogonal, but they are robust when LD between two SNPs is low, allowing the fast screening achieved by BiForce. When LD is high, however, the test for interaction between two correlated SNPs is inflated by the marginal effects of the pair of SNPs, and therefore the inflation is critical when marginal effects are strong but not when marginal effects are weak. Nevertheless, because BiForce uses stringent Bonferroni-adjusted thresholds by default, the chance of inflated epistatic pairs being genomewide significant should be low in general. This issue affected the results in Table 2 of the original article (Gyenesei et al., 2012). The correct interaction P-values (Pint) of each SNP pair are listed in the updated Table 2 later in the text, suggesting that none remained genome-wide significant in C-reactive protein (CRP), glucose (GLU), high-density lipoprotein (HDL), low-density lipoprotein (LDL) and triglycerides (TRI). Results from the analyses of simulated data on quantitative and disease traits are unaffected because in simulation SNPs were randomly drawn from a chromosome assuming they were in Hardy–Weinberg equilibrium, ie the chance of high LD coming together with strong marginal effects is very low, which is supported by the results of false positive rate in Figure 3 of the original article. In summary, the issue only affects a small part of