The sum-over-forests density index: Identifying dense regions in a graph

The sum-over-forests density index: Identifying dense regions in a graph
复制标题

森林总和密度指数:识别图中的密集区域

DOI:
10.1109/tpami.2013.227
复制
发表时间:
2014
影响因子:
23.6
通讯作者:
Francois Fouss
Francois Fouss
中科院分区:
计算机科学1区
文献类型:
--
作者:
Mathieu Senelle;Silvia Garcia-Diez;Amin Mantrach;Masashi Shimbo;Marco Saerens;Francois Fouss

文献摘要

相似文献

本文介绍了一种定义在图上的非参数密度指数,即森林和密度指数。它基于一个清晰直观的思想:图中的高密度区域的特征是包含大量高外度的低成本树,而低密度区域包含很少的树。因此,定义图中可数森林集合上的玻尔兹曼概率分布,使得大(高成本)森林以低概率出现,而短(低成本)森林以高概率出现。然后,将节点的sofm密度指数定义为该节点在森林集合上的期望超出度,从而提供该节点周围密度的度量。根据矩阵森林定理和统计物理框架,通过简单的矩阵反演,可以很容易地以封闭形式计算软密度指数。在人工数据集和真实数据集上的实验表明,对于不同来源的图,所提出的索引在寻找密集区域方面表现良好。
This work introduces a novel nonparametric density index defined on graphs, the Sum-over-Forests (SoF) density index. It is based on a clear and intuitive idea: high-density regions in a graph are characterized by the fact that they contain a large amount of low-cost trees with high outdegrees while low-density regions contain few ones. Therefore, a Boltzmann probability distribution on the countable set of forests in the graph is defined so that large (high-cost) forests occur with a low probability while short (low-cost) forests occur with a high probability. Then, the SoF density index of a node is defined as the expected outdegree of this node on the set of forests, thus providing a measure of density around that node. Following the matrix-forest theorem and a statistical physics framework, it is shown that the SoF density index can be easily computed in closed form through a simple matrix inversion. Experiments on artificial and real datasets show that the proposed index performs well on finding dense regions, for graphs of various origins.