Graph-based Security and Privacy Analytics via Collective Classification with Joint Weight Learning and Propagation

Graph-based Security and Privacy Analytics via Collective Classification with Joint Weight Learning and Propagation
复制标题

DOI:
10.14722/ndss.2019.23226
复制
发表时间:
2018-12
期刊:
ArXiv
影响因子:
--
通讯作者:
Binghui Wang;Jinyuan Jia;N. Gong
Binghui Wang;Jinyuan Jia;N. Gong
中科院分区:
其他
文献类型:
--
作者:
Binghui Wang;Jinyuan Jia;N. Gong

文献摘要

被引文献

相似文献

许多安全和隐私问题可以建模为一个图分类问题,其中图中的节点同时通过集体分类进行分类。这种基于图的安全和隐私分析的最先进的集体分类方法遵循以下范式:为图的边缘分配权重,迭代地在加权图中传播节点的信誉分数,并使用最终的信誉分数对图中的节点进行分类。关键的挑战是分配边缘权重,如果两个对应的节点具有相同的标签,则边缘具有较大的权重,否则则具有较小的权重。尽管集体分类在安全和隐私问题上的研究和应用已有十多年,但如何解决这一挑战仍然是一个悬而未决的问题。在这项工作中,我们提出了一个新的集体分类框架来解决这个长期存在的挑战。我们首先将学习边缘权重作为一个优化问题,它量化了我们要达到的最终声誉分数的目标。然而,由于最终声誉分数以一种非常复杂的方式依赖于边缘权重,因此在计算上很难解决优化问题。为了解决计算挑战,我们提出联合学习边缘权重和传播声誉分数,这本质上是优化问题的近似解决方案。我们将我们的框架与最先进的基于图形的安全和隐私分析方法进行比较,使用来自各种应用场景的四个大规模真实世界数据集,如社交网络中的Sybil检测,Yelp中的虚假评论检测和属性推理攻击。我们的结果表明,我们的框架在可接受的计算开销下实现了比最先进的方法更高的精度。
Many security and privacy problems can be modeled as a graph classification problem, where nodes in the graph are classified by collective classification simultaneously. State-of-the-art collective classification methods for such graph-based security and privacy analytics follow the following paradigm: assign weights to edges of the graph, iteratively propagate reputation scores of nodes among the weighted graph, and use the final reputation scores to classify nodes in the graph. The key challenge is to assign edge weights such that an edge has a large weight if the two corresponding nodes have the same label, and a small weight otherwise. Although collective classification has been studied and applied for security and privacy problems for more than a decade, how to address this challenge is still an open question. In this work, we propose a novel collective classification framework to address this long-standing challenge. We first formulate learning edge weights as an optimization problem, which quantifies the goals about the final reputation scores that we aim to achieve. However, it is computationally hard to solve the optimization problem because the final reputation scores depend on the edge weights in a very complex way. To address the computational challenge, we propose to jointly learn the edge weights and propagate the reputation scores, which is essentially an approximate solution to the optimization problem. We compare our framework with state-of-the-art methods for graph-based security and privacy analytics using four large-scale real-world datasets from various application scenarios such as Sybil detection in social networks, fake review detection in Yelp, and attribute inference attacks. Our results demonstrate that our framework achieves higher accuracies than state-of-the-art methods with an acceptable computational overhead.