Estimating Aggregate Properties In Relational Networks With Unobserved Data

Estimating Aggregate Properties In Relational Networks With Unobserved Data
复制标题

DOI:
--
复制
发表时间:
2020-01
期刊:
ArXiv
影响因子:
--
通讯作者:
Varun R. Embar;S. Srinivasan;L. Getoor
Varun R. Embar;S. Srinivasan;L. Getoor
中科院分区:
其他
文献类型:
--
作者:
Varun R. Embar;S. Srinivasan;L. Getoor

文献摘要

相似文献

群集内聚性和桥接节点数量等聚合网络属性可用于收集有关网络社区结构、影响力传播和网络对故障的恢复能力的见解。在充分观察网络的情况下,高效地计算网络属性受到了极大的关注(Wasserman and Faust 1994; Cook and Holder 2006),然而,在数据属性缺失的情况下,计算聚合网络属性的问题却很少受到关注。为缺少属性的网络计算这些属性需要在网络上执行推理。统计关系学习(SRL)和图神经网络(gnn)是两类非常适合推断图中缺失属性的机器学习方法。在本文中,我们研究了这些方法在估计缺失属性网络的聚合属性方面的有效性。我们比较了两种SRL方法和三种gnn。对于这些方法,我们使用点估计(如MAP和mean)来估计这些属性。对于可以推断缺失属性的联合分布的基于srl的方法,我们也将这些属性作为对分布的期望来估计。为了可跟踪地计算概率软逻辑(我们研究的SRL方法之一)的期望,我们引入了一个新的采样框架。在使用三个基准数据集的实验评估中,我们表明基于srl的方法在计算聚合属性和预测精度方面都优于基于gnn的方法。具体地说,我们表明,估计总体属性作为联合分布的期望优于点估计。
Aggregate network properties such as cluster cohesion and the number of bridge nodes can be used to glean insights about a network's community structure, spread of influence and the resilience of the network to faults. Efficiently computing network properties when the network is fully observed has received significant attention (Wasserman and Faust 1994; Cook and Holder 2006), however the problem of computing aggregate network properties when there is missing data attributes has received little attention. Computing these properties for networks with missing attributes involves performing inference over the network. Statistical relational learning (SRL) and graph neural networks (GNNs) are two classes of machine learning approaches well suited for inferring missing attributes in a graph. In this paper, we study the effectiveness of these approaches in estimating aggregate properties on networks with missing attributes. We compare two SRL approaches and three GNNs. For these approaches we estimate these properties using point estimates such as MAP and mean. For SRL-based approaches that can infer a joint distribution over the missing attributes, we also estimate these properties as an expectation over the distribution. To compute the expectation tractably for probabilistic soft logic, one of the SRL approaches that we study, we introduce a novel sampling framework. In the experimental evaluation, using three benchmark datasets, we show that SRL-based approaches tend to outperform GNN-based approaches both in computing aggregate properties and predictive accuracy. Specifically, we show that estimating the aggregate properties as an expectation over the joint distribution outperforms point estimates.