Resampling-based predictive simulation framework of stochastic diffusion model for identifying top-K influential nodes

Resampling-based predictive simulation framework of stochastic diffusion model for identifying top-K influential nodes
复制标题

DOI:
10.1007/s41060-019-00183-3
复制
发表时间:
2013-10
影响因子:
2.4
通讯作者:
K. Ohara;Kazumi Saito;M. Kimura;H. Motoda
K. Ohara;Kazumi Saito;M. Kimura;H. Motoda
中科院分区:
--
文献类型:
--
作者:
K. Ohara;Kazumi Saito;M. Kimura;H. Motoda

文献摘要

相似文献

我们解决了有效估计节点在社交网络上信息传播中的影响的问题。由于信息扩散是一个随机过程,节点的影响程度是通过期望来量化的,而期望通常是通过非常耗时的多次模拟来获得的。我们的贡献是,我们提出了一种基于留置交叉验证技术的预测模拟框架,该框架很好地近似了两个目标问题的未知地面事实的误差:一个用于估计每个节点的影响程度,另一个用于识别顶级关联节点。我们针对第一个问题提出的方法估计每个节点的影响度的近似误差,而针对第二个问题的方法估计导出的top-Knodes的精度,两者都在不知道真实影响度的情况下。我们使用三个现实世界网络对所提出的方法进行实验评估,结果表明,当使用留半交叉验证(即 N 是运行次数的一半)时,它们可以作为一种很好的方法来解决目标问题,并且在使用留半交叉验证(即 Ni 是运行次数的一半)时,可以用更少的模拟运行来解决目标问题,这意味着人们可以在不确切知道影响程度的情况下以良好的精度识别有影响的节点。
We address a problem of efficiently estimating the influence of a node in information diffusion over a social network. Since the information diffusion is a stochastic process, the influence degree of a node is quantified by the expectation, which is usually obtained by very time-consuming many runs of simulation. Our contribution is that we proposed a framework for predictive simulation based on the leave-N-out cross-validation technique that well approximates the error from the unknown ground truth for two target problems: one to estimate the influence degree of each node, and the other to identify top-Kinfluential nodes. The method we proposed for the first problem estimates the approximation error of the influence degree of each node, and the method for the second problem estimates the precision of the derived top-Knodes, both without knowing the true influence degree. We experimentally evaluate the proposed methods using the three real-world networks and show that they can serve as a good measure to solve the target problems with far fewer runs of simulation ensuring the accuracy when the leave-half-out cross-validation, i.e.,Nis the half of the number of runs, is used, which means that one can identify the influential nodes without knowing exactly their influence degree in good accuracy.