Identification of genetic interaction networks via an evolutionary algorithm evolved Bayesian network.

Identification of genetic interaction networks via an evolutionary algorithm evolved Bayesian network.
复制标题

DOI:
10.1186/s13040-016-0094-4
复制
发表时间:
2016
期刊:
影响因子:
4.5
通讯作者:
Ritchie MD
Ritchie MD
中科院分区:
生物学3区
文献类型:
--
作者:
Li R;Dudek SM;Kim D;Hall MA;Bradford Y;Peissig PL;Brilliant MH;Linneman JG;McCarty CA;Bao L;Ritchie MD

文献摘要

被引文献

相似文献

医学的未来正在走向精准医学阶段,目标是通过考虑个体间的差异性来预防和治疗疾病。这种可变性很大程度上取决于我们的基因构成。随着高通量基因组测序方法的快速发展,已经产生了海量的遗传学数据。精准医学的下一个障碍是拥有足够的计算工具来分析大量数据。全基因组关联研究已成为评价单核苷酸多态(SNPs)与疾病性状关系的主要方法。虽然GWAS足以发现具有较强主效应的单个SNPs,但它不能捕捉到多个SNPs之间的潜在相互作用。在许多性状中,很大一部分变异仍然不能仅用主效应来解释,这为探索遗传交互作用的作用敞开了大门。然而,在大规模基因组数据中识别基因相互作用即使对现代计算也是一个挑战。在这项研究中,我们提出了一种新的算法,语法进化贝叶斯网络(GEBN),它利用贝叶斯网络来识别数据中的交互,同时使用进化算法来降低与网络优化相关的计算成本。GEBN在数据包含主效应和交互效应的模拟研究中表现出色。我们还将GEBN应用于从Marshfield个性化医学研究项目(PMRP)获得的2型糖尿病(T2D)数据集。我们能够确定T2D病例和对照的遗传交互作用,并使用来自这些交互作用的信息来对T2D样本进行分类。我们得到了平均曲线下面积(AUC)为86.8%。我们还鉴定了几个相互作用的基因,如INADL和LPP,这些基因已知与T2D相关。开发计算机工具来探索主效应以外的遗传关联仍然是人类遗传学中的一个至关重要的挑战。像GEBN这样的方法展示了考虑遗传交互作用的效用,因为它们可能解释了一些缺失的遗传性。
The future of medicine is moving towards the phase of precision medicine, with the goal to prevent and treat diseases by taking inter-individual variability into account. A large part of the variability lies in our genetic makeup. With the fast paced improvement of high-throughput methods for genome sequencing, a tremendous amount of genetics data have already been generated. The next hurdle for precision medicine is to have sufficient computational tools for analyzing large sets of data. Genome-Wide Association Studies (GWAS) have been the primary method to assess the relationship between single nucleotide polymorphisms (SNPs) and disease traits. While GWAS is sufficient in finding individual SNPs with strong main effects, it does not capture potential interactions among multiple SNPs. In many traits, a large proportion of variation remain unexplained by using main effects alone, leaving the door open for exploring the role of genetic interactions. However, identifying genetic interactions in large-scale genomics data poses a challenge even for modern computing. For this study, we present a new algorithm, Grammatical Evolution Bayesian Network (GEBN) that utilizes Bayesian Networks to identify interactions in the data, and at the same time, uses an evolutionary algorithm to reduce the computational cost associated with network optimization. GEBN excelled in simulation studies where the data contained main effects and interaction effects. We also applied GEBN to a Type 2 diabetes (T2D) dataset obtained from the Marshfield Personalized Medicine Research Project (PMRP). We were able to identify genetic interactions for T2D cases and controls and use information from those interactions to classify T2D samples. We obtained an average testing area under the curve (AUC) of 86.8 %. We also identified several interacting genes such as INADL and LPP that are known to be associated with T2D. Developing the computational tools to explore genetic associations beyond main effects remains a critically important challenge in human genetics. Methods, such as GEBN, demonstrate the utility of considering genetic interactions, as they likely explain some of the missing heritability.