Model averaging strategies for structure learning in Bayesian networks with limited data.

Model averaging strategies for structure learning in Bayesian networks with limited data.
复制标题

DOI:
10.1186/1471-2105-13-s13-s10
复制
发表时间:
2012
期刊:
影响因子:
3
通讯作者:
Subramanian D
Subramanian D
中科院分区:
生物学4区
文献类型:
--
作者:
Broom BM;Do KA;Subramanian D

文献摘要

被引文献

相似文献

从数据中学习贝叶斯网络结构的算法已经取得了相当大的进展。通过使用自举复制与通过阈值化的特征选择的模型平均是用于学习具有高置信度的特征的广泛使用的解决方案。然而,由于数据有限,许多问题仍然没有答案。什么评分函数对于模型平均最有效?自举法的离散性是否会显著影响学习绩效?选择单个最佳网络还是平均从每个bootstrap重采样中学习的多个网络更好?如何选择学习统计显著特征的阈值?最好的评分函数是小λ的Dirichlet先验评分度量和贝叶斯Dirichlet度量。修正了自举算法的离散性所引起的偏差,提高了自举算法的学习性能。最好选择从每个bootstrap重采样中学习到的单个最佳网络。我们描述了一种基于置换的方法,用于确定在袋装模型中进行特征选择的显著性阈值。我们表明,在有限的数据的情况下,贝叶斯装袋使用狄利克雷先验评分度量(DPSM)是最有效的学习策略,修改评分函数,惩罚复杂的网络阻碍模型平均。我们使用两个著名的基准,特别是报警和保险的系统研究建立这些结果。我们还将我们的网络构建方法应用于来自癌症基因组图谱多形性胶质母细胞瘤数据集的基因表达数据,并表明生存率与临床协变量年龄和性别以及干扰素诱导基因和生长抑制基因的簇有关。对于小数据集,我们的方法比以前发表的方法表现得更好。
Considerable progress has been made on algorithms for learning the structure of Bayesian networks from data. Model averaging by using bootstrap replicates with feature selection by thresholding is a widely used solution for learning features with high confidence. Yet, in the context of limited data many questions remain unanswered. What scoring functions are most effective for model averaging? Does the bias arising from the discreteness of the bootstrap significantly affect learning performance? Is it better to pick the single best network or to average multiple networks learnt from each bootstrap resample? How should thresholds for learning statistically significant features be selected? The best scoring functions are Dirichlet Prior Scoring Metric with small λ and the Bayesian Dirichlet metric. Correcting the bias arising from the discreteness of the bootstrap worsens learning performance. It is better to pick the single best network learnt from each bootstrap resample. We describe a permutation based method for determining significance thresholds for feature selection in bagged models. We show that in contexts with limited data, Bayesian bagging using the Dirichlet Prior Scoring Metric (DPSM) is the most effective learning strategy, and that modifying the scoring function to penalize complex networks hampers model averaging. We establish these results using a systematic study of two well-known benchmarks, specifically ALARM and INSURANCE. We also apply our network construction method to gene expression data from the Cancer Genome Atlas Glioblastoma multiforme dataset and show that survival is related to clinical covariates age and gender and clusters for interferon induced genes and growth inhibition genes. For small data sets, our approach performs significantly better than previously published methods.