High-dimensional structure learning of binary pairwise Markov networks: A comparative numerical study

High-dimensional structure learning of binary pairwise Markov networks: A comparative numerical study
复制标题

二元成对马尔可夫网络的高维结构学习:比较数值研究

DOI:
10.1016/j.csda.2019.06.012
复制
发表时间:
2020
影响因子:
1.8
通讯作者:
Corander Jukka
Corander Jukka
中科院分区:
数学3区
文献类型:
--
作者:
Pensar Johan;Xu Yingying;Puranen Santeri;Pesonen Maiju;Kabashima Yoshiyuki;Corander Jukka

文献摘要

相似文献

从数据中学习马尔可夫网络的无向图结构是一个在过去几十年中受到广泛关注的问题。由于模型类的普遍适用性,在几个研究领域中并行开发了无数的方法。最近,随着所考虑的系统的大小已经增加,新方法的重点已经转移到高维域。特别是,伪似然函数的引入已经推动了最初基于似然函数的基于分数的方法的极限。与此同时,基于简单成对检验的方法已经被开发出来,以满足计算生物学中越来越大的数据集所带来的挑战。除了适用于高维问题,基于伪似然和成对检验的方法从根本上是非常不同的。为了比较不同类型的方法的准确性,进行了广泛的数值研究所产生的二进制成对马尔可夫网络的数据。提出了一种基于受限玻尔兹曼机的可并行化Gibbs采样器,作为从稀疏高维网络中高效采样的工具。研究结果表明,在高维结构学习应用中经常遇到的设置中,成对方法可以比伪似然方法更准确。
Learning the undirected graph structure of a Markov network from data is a problem that has received a lot of attention during the last few decades. As a result of the general applicability of the model class, a myriad of methods have been developed in parallel in several research fields. Recently, as the size of the considered systems has increased, the focus of new methods has been shifted towards the high-dimensional domain. In particular, introduction of the pseudo-likelihood function has pushed the limits of score-based methods which were originally based on the likelihood function. At the same time, methods based on simple pairwise tests have been developed to meet the challenges arising from increasingly large data sets in computational biology. Apart from being applicable to high-dimensional problems, methods based on the pseudo-likelihood and pairwise tests are fundamentally very different. To compare the accuracy of the different types of methods, an extensive numerical study is performed on data generated by binary pairwise Markov networks. A parallelizable Gibbs sampler, based on restricted Boltzmann machines, is proposed as a tool to efficiently sample from sparse high-dimensional networks. The results of the study show that pairwise methods can be more accurate than pseudo-likelihood methods in settings often encountered in high-dimensional structure learning applications.