BOND: Benchmarking Unsupervised Outlier Node Detection on Static Attributed Graphs

BOND: Benchmarking Unsupervised Outlier Node Detection on Static Attributed Graphs
复制标题

DOI:
--
复制
发表时间:
2022-06
期刊:
--
影响因子:
--
通讯作者:
Kay Liu;Yingtong Dou;Yue Zhao;Xueying Ding;Xiyang Hu;Ruitong Zhang;Kaize Ding;Canyu Chen;Hao Peng;Kai Shu;Lichao Sun;Jundong Li;George H. Chen;Zhihao Jia;Philip S. Yu
Kay Liu;Yingtong Dou;Yue Zhao;Xueying Ding;Xiyang Hu;Ruitong Zhang;Kaize Ding;Canyu Chen;Hao Peng;Kai Shu;Lichao Sun;Jundong Li;George H. Chen;Zhihao Jia;Philip S. Yu
中科院分区:
其他
文献类型:
--
作者:
Kay Liu;Yingtong Dou;Yue Zhao;Xueying Ding;Xiyang Hu;Ruitong Zhang;Kaize Ding;Canyu Chen;Hao Peng;Kai Shu;Lichao Sun;Jundong Li;George H. Chen;Zhihao Jia;Philip S. Yu

文献摘要

相似文献

检测图中哪些节点是离群值是一个相对较新的机器学习任务,有许多应用。尽管近年来针对这一任务开发了大量算法,但还没有一个标准的综合性能评估设置。因此,很难理解哪些方法在广泛的设置下工作良好。为了弥补这一差距,据我们所知,我们提出了静态属性图上无监督离群节点检测的第一个综合基准,称为BOND,具有以下亮点。(1)我们对从经典的矩阵分解到最新的图神经网络的14种方法的异常点检测性能进行了基准测试。(2)使用9个真实数据集,我们的基准评估了不同的检测方法对两种主要类型的合成异常值的响应,以及分别对“有机”(真正的非合成)异常值的响应。(3)利用现有的随机图生成技术,我们生成了一系列不同图大小的综合生成数据集,使我们能够比较不同离群点检测算法的运行时间和内存使用情况。基于我们的实验结果,我们讨论了现有图离群点检测算法的优缺点,并强调了未来研究的机会。重要的是,我们的代码是免费提供的,并且易于扩展:https://github.com/pygod-team/pygod/tree/main/benchmark
Detecting which nodes in graphs are outliers is a relatively new machine learning task with numerous applications. Despite the proliferation of algorithms developed in recent years for this task, there has been no standard comprehensive setting for performance evaluation. Consequently, it has been difficult to understand which methods work well and when under a broad range of settings. To bridge this gap, we present--to the best of our knowledge--the first comprehensive benchmark for unsupervised outlier node detection on static attributed graphs called BOND, with the following highlights. (1) We benchmark the outlier detection performance of 14 methods ranging from classical matrix factorization to the latest graph neural networks. (2) Using nine real datasets, our benchmark assesses how the different detection methods respond to two major types of synthetic outliers and separately to"organic"(real non-synthetic) outliers. (3) Using an existing random graph generation technique, we produce a family of synthetically generated datasets of different graph sizes that enable us to compare the running time and memory usage of the different outlier detection algorithms. Based on our experimental results, we discuss the pros and cons of existing graph outlier detection algorithms, and we highlight opportunities for future research. Importantly, our code is freely available and meant to be easily extendable: https://github.com/pygod-team/pygod/tree/main/benchmark