Learning Predictive Interactions Using Information Gain and Bayesian Network Scoring.

Learning Predictive Interactions Using Information Gain and Bayesian Network Scoring.
复制标题

DOI:
10.1371/journal.pone.0143247
复制
发表时间:
2015
期刊:
影响因子:
3.7
通讯作者:
Neapolitan R
Neapolitan R
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Jiang X;Jao J;Neapolitan R

文献摘要

被引文献

相似文献

在统计和机器学习领域中,相关性和分类问题一直存在,并且已经开发了解决这些问题的技术。我们现在处于高维数据时代,高维数据可以涉及数十亿个变量。这些数据提出了新的挑战。特别是,当每个变量的边际效应很小时,很难发现预测变量。一个例子涉及全基因组关联研究(GWAS)数据集,其中涉及数百万个单核苷酸多态性(snp),其中一些snp通过上位性相互作用影响疾病状态。为了确定这些相互作用的snp,研究人员开发了解决这一特定问题的技术。然而,这个问题更为普遍,因此这些技术适用于与交互有关的其他问题。许多这些技术的一个困难是,它们不能区分一个习得的相互作用是否真的是一个相互作用,或者它是否涉及几个具有强烈边际效应的变量。我们使用信息增益和贝叶斯网络评分来解决这个问题。首先,我们通过确定变量是否一起提供比单独提供更多的信息来确定候选交互。然后,我们使用贝叶斯网络评分来查看候选交互是否真的是一个可能的模型。我们的策略被称为mbs -再融资。使用100个模拟数据集和一个真实的GWAS阿尔茨海默病数据集,我们研究了MBS-IGain的性能。在分析模拟数据集时,MBS-IGain在定位相互作用预测因子和准确识别相互作用方面大大优于之前的九种方法。在分析真实的阿尔茨海默病数据集时,我们获得了新的结果,并且结果证实了之前的发现。我们得出结论,MBS-IGain在寻找高维数据集中的相互作用方面非常有效。这个结果很重要,因为我们在许多领域拥有越来越丰富的高维数据,为了学习原因并使用这些数据进行预测/分类,我们通常必须首先确定相互作用。
The problems of correlation and classification are long-standing in the fields of statistics and machine learning, and techniques have been developed to address these problems. We are now in the era of high-dimensional data, which is data that can concern billions of variables. These data present new challenges. In particular, it is difficult to discover predictive variables, when each variable has little marginal effect. An example concerns Genome-wide Association Studies (GWAS) datasets, which involve millions of single nucleotide polymorphism (SNPs), where some of the SNPs interact epistatically to affect disease status. Towards determining these interacting SNPs, researchers developed techniques that addressed this specific problem. However, the problem is more general, and so these techniques are applicable to other problems concerning interactions. A difficulty with many of these techniques is that they do not distinguish whether a learned interaction is actually an interaction or whether it involves several variables with strong marginal effects. We address this problem using information gain and Bayesian network scoring. First, we identify candidate interactions by determining whether together variables provide more information than they do separately. Then we use Bayesian network scoring to see if a candidate interaction really is a likely model. Our strategy is called MBS-IGain. Using 100 simulated datasets and a real GWAS Alzheimer’s dataset, we investigated the performance of MBS-IGain. When analyzing the simulated datasets, MBS-IGain substantially out-performed nine previous methods at locating interacting predictors, and at identifying interactions exactly. When analyzing the real Alzheimer’s dataset, we obtained new results and results that substantiated previous findings. We conclude that MBS-IGain is highly effective at finding interactions in high-dimensional datasets. This result is significant because we have increasingly abundant high-dimensional data in many domains, and to learn causes and perform prediction/classification using these data, we often must first identify interactions.