ImVerde: Vertex-Diminished Random Walk for Learning Imbalanced Network Representation

ImVerde: Vertex-Diminished Random Walk for Learning Imbalanced Network Representation
复制标题

DOI:
10.1109/bigdata.2018.8622603
复制
发表时间:
2018-04
期刊:
2018 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Jun Wu;Jingrui He;Yongming Liu
Jun Wu;Jingrui He;Yongming Liu
中科院分区:
其他
文献类型:
--
作者:
Jun Wu;Jingrui He;Yongming Liu

文献摘要

相似文献

不平衡数据广泛存在于许多高影响力的应用中。例如,在空中交通管制中,在所有三种类型的事故原因中,具有“人员问题”的历史事故报告比其他两种类型(“飞机问题”和“环境问题”)的总和要多得多。因此,由此产生的事故报告数据集是高度不平衡的。另一方面,该数据集可以自然地建模为网络,其中每个节点表示事故报告,并且每个边指示一对事故报告的相似性。到目前为止,大多数不平衡数据分析的工作集中在分类设置,很少致力于学习节点表示的不平衡网络。为了弥补这一差距,在本文中,我们首先提出顶点缩减随机游走(VDRW)的不平衡网络分析。它与现有的顶点增强随机行走有很大的不同,因为它不鼓励随机粒子返回到已经访问过的节点。这种设计特别适用于不平衡网络,因为随机粒子更有可能访问同一类的节点,这是学习节点表示的理想属性。在此基础上,提出了一种基于VDRW的半监督网络表示学习框架ImVerde,其中上下文采样使用VDRW和有限的标签信息来创建节点-上下文对,而平衡批量采样采用一种简单的欠采样方法来平衡来自不同类的节点-上下文对.实验结果表明,基于VDRW的ImVerde在从不平衡数据中学习网络表示方面优于最先进的算法。
Imbalanced data widely exist in many high-impact applications. An example is in air traffic control, where among all three types of accident causes, historical accident reports with ‘personnel issues’ are much more than the other two types (‘aircraft issues’ and ‘environmental issues’) combined. Thus, the resulting data set of accident reports is highly imbalanced. On the other hand, this data set can be naturally modeled as a network, with each node representing an accident report, and each edge indicating the similarity of a pair of accident reports. Up until now, most existing work on imbalanced data analysis focused on the classification setting, and very little is devoted to learning the node representations for imbalanced networks. To bridge this gap, in this paper, we first propose Vertex-Diminished Random Walk (VDRW) for imbalanced network analysis. It is significantly different from the existing Vertex Reinforced Random Walk by discouraging the random particle to return to the nodes that have already been visited. This design is particularly suitable for imbalanced networks as the random particle is more likely to visit the nodes from the same class, which is a desired property for learning node representations. Furthermore, based on VDRW, we propose a semi-supervised network representation learning framework named ImVerde for imbalanced networks, where context sampling uses VDRW and the limited label information to create node-context pairs, and balanced-batch sampling adopts a simple under-sampling method to balance these pairs from different classes. Experimental results demonstrate that ImVerde based on VDRW outperforms state-of-the-art algorithms for learning network representations from imbalanced data.