Entropy-based information gain approaches to detect and to characterize gene-gene and gene-environment interactions/correlations of complex diseases.

Entropy-based information gain approaches to detect and to characterize gene-gene and gene-environment interactions/correlations of complex diseases.
复制标题

DOI:
10.1002/gepi.20621
复制
发表时间:
2011-11
影响因子:
2.1
通讯作者:
Moore, J. H.
Moore, J. H.
中科院分区:
医学4区
文献类型:
--
作者:
Fan, R.;Zhong, M.;Wang, S.;Zhang, Y.;Andrew, A.;Karagas, M.;Chen, H.;Amos, C. I.;Xiong, M.;Moore, J. H.

文献摘要

参考文献

被引文献

相似文献

对于复杂的疾病,基因分型、环境因素和表型之间的关系通常是复杂的和非线性的。在过去的几年里,我们对疾病的遗传结构的理解有了很大的提高。然而,尽管存在许多有效的方法,但在概念和方法上,检测基因-基因和基因-环境相互作用仍然是一个挑战。有一种方法前景看好,但尚未广泛应用于基因组数据,那就是信息论的基于熵的方法。在本文中,我们首先发展了基于熵的检验统计量来识别双向和更高阶的基因-基因和基因-环境相互作用。然后,我们将这些方法应用于膀胱癌数据集,从而测试它们的能力并确定其优势和劣势。对于双向交互,我们提出了一种基于互信息的信息增益方法。对于三次和更高阶的相互作用,采用交互-信息-增益方法。在这两种情况下,我们都开发了一维测试统计来分析稀疏数据。与朴素卡方检验相比,我们开发的检验统计量具有类似或更高的能力,并且是稳健的。将其应用于膀胱癌数据集,可以研究DNA修复基因SNPs、吸烟状况和膀胱癌易感性之间的复杂交互作用。虽然尚未广泛应用,但基于熵的方法似乎是检测基因-基因和基因-环境相互作用的有用工具。我们开发的测试统计数据增加了越来越多的身体方法,这些方法将逐渐揭示常见疾病的复杂架构。
For complex diseases, the relationship between genotypes, environment factors and phenotype is usually complex and nonlinear. Our understanding of the genetic architecture of diseases has considerably increased over the last years. However, both conceptually and methodologically, detecting gene-gene and gene-environment interactions remains a challenge, despite the existence of a number of efficient methods. One method that offers great promises but has not yet been widely applied to genomic data is the entropy-based approach of information theory. In this paper we first develop entropy-based test statistics to identify 2-way and higher order gene-gene and gene-environment interactions. We then apply these methods to a bladder cancer data set and thereby test their power and identify strengths and weaknesses. For two-way interactions, we propose an information-gain approach based on mutual information. For three-ways and higher order interactions, an interaction-information-gain approach is used. In both case we develop one-dimensional test statistics to analyze sparse data. Compared to the naive chi-square test, the test statistics we develop have similar or higher power and is robust. Applying it to the bladder cancer data set allowed to investigate the complex interactions between DNA repair gene SNPs, smoking status, and bladder cancer susceptibility. Although not yet widely applied, entropy-based approaches appear as a useful tool for detecting gene-gene and gene-environment interactions. The test statistics we develop add to a growing body methodologies that will gradually shed light on the complex architecture of common diseases.
DOI: 10.1002/j.1538-7305.1948.tb01338.x
发表时间: 1948-01-01
影响因子: --
作者:
SHANNON, CE
通讯作者: SHANNON, CE
DOI: 10.1002/gepi.20128
发表时间: 2006-02-01
影响因子: 2.1
作者:
Martin, ER;Ritchie, MD;Moore, JH
通讯作者: Moore, JH
DOI: 10.1186/1471-2105-4-28
发表时间: 2003-07-07
期刊: BMC bioinformatics
影响因子: 3
作者:
Ritchie MD;White BC;Parker JS;Hahn LW;Moore JH
通讯作者: Moore JH
DOI: 10.1086/321276
发表时间: 2001-07-01
影响因子: 9.8
作者:
Ritchie, MD;Hahn, LW;Moore, JH
通讯作者: Moore, JH
DOI: 10.1093/bioinformatics/btf869
发表时间: 2003-02-12
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Hahn, LW;Ritchie, MD;Moore, JH
通讯作者: Moore, JH