Lime: Low-Cost and Incremental Learning for Dynamic Heterogeneous Information Networks

Lime: Low-Cost and Incremental Learning for Dynamic Heterogeneous Information Networks
复制标题

DOI:
10.1109/tc.2021.3057082
复制
发表时间:
2021-02
影响因子:
3.7
通讯作者:
Hao Peng-;Renyu Yang;Z. Wang;Jianxin Li;Lifang He;Philip S. Yu;A. Zomaya;R. Ranjan
Hao Peng-;Renyu Yang;Z. Wang;Jianxin Li;Lifang He;Philip S. Yu;A. Zomaya;R. Ranjan
中科院分区:
计算机科学2区
文献类型:
--
作者:
Hao Peng-;Renyu Yang;Z. Wang;Jianxin Li;Lifang He;Philip S. Yu;A. Zomaya;R. Ranjan

文献摘要

被引文献

相似文献

了解社交网络、学者网络和物联网网络等大规模信息网络的相互关联关系,对于推荐和欺诈检测等任务至关重要。绝大多数现实世界的网络本质上是异质的和动态的,包含许多不同类型的节点和边,并且可能会随着时间的推移而发生巨大变化。这种动态性和异构性使得对网络结构的推理具有极大的挑战性。遗憾的是,现有的方法要么对给定的随机过程有很强的假设,要么不能捕捉到网络结构的异质性,都需要大量的计算资源,不能很好地对现实生活中的动态网络进行建模。我们介绍了Lime,这是一种更好地建模动态和异质信息网络的方法。LIME旨在提取高质量的网络表示,而内存资源和计算时间比最先进的要低得多。与以往使用向量对每个网络节点进行编码不同,我们利用网络节点之间的语义关系来对共享向量中具有相似语义的多个节点进行编码。通过使用更少的节点向量,我们的方法显著减少了编码大规模网络所需的存储空间。为了有效地用信息共享来换取更少的存储空间,我们使用了递归神经网络(RsNN)和精心设计的优化策略来探索一个新的长方体空间中的节点语义。然后,我们进一步展示了如何在RsNN、我们的长方体结构和一组新的优化技术的帮助下,开发出一种有效的增量学习方法,以允许学习框架快速高效地适应不断发展的网络。我们通过将其应用于三个典型的基于网络的任务:节点分类、节点聚类和异常检测来评估LIME,并在三个大规模数据集上执行。我们将Lime与11种学习网络表示的先前最先进的方法进行了比较。我们的大量实验表明,在学习网络表示时,Lime不仅将内存占用减少了80%以上,处理时间减少了2倍以上,而且对于下游处理任务也提供了类似的性能。我们的增量学习方法可以在不影响学习网络表示质量的情况下将学习时间提高到原来的20倍。
Understanding the interconnected relationships of large-scale information networks like social, scholar and Internet of Things networks is vital for tasks like recommendation and fraud detection. The vast majority of the real-world networks are inherently heterogeneous and dynamic, containing many different types of nodes and edges and can change drastically over time. The dynamicity and heterogeneity make it extremely challenging to reason about the network structure. Unfortunately, existing approaches are inadequate in modeling real-life dynamical networks as they either have strong assumption of a given stochastic process or fail to capture the heterogeneity of network structure, and they all require extensive computational resources. We introduce Lime, a better approach for modeling dynamic and heterogeneous information networks. Lime is designed to extract high-quality network representation with significantly lower memory resources and computational time over the state-of-the-arts. Unlike prior work that uses a vector to encode each network node, we exploit the semantic relationships among network nodes to encode multiple nodes with similar semantics in shared vectors. By using many fewer node vectors, our approach significantly reduces the required memory space for encoding large-scale networks. To effectively trade information sharing for reduced memory footprint, we employ the recursive neural network (RsNN) with carefully designed optimization strategies to explore the node semantics in a novel cuboid space. We then go further by showing, for the first time, how an effective incremental learning approach can be developed – with the help of RsNN, our cuboid structure, and a set of novel optimization techniques – to allow a learning framework to quickly and efficiently adapt to a constantly evolving network. We evaluate Lime by applying it to three representative network-based tasks, node classification, node clustering and anomaly detection, performing on three large-scale datasets. We compare Lime against eleven prior state-of-the-art approaches for learning network representation. Our extensive experiments demonstrate that Lime not only reduces the memory footprint by over 80 percent and the processing time over 2x when learning network representation but also delivers comparable performance for downstream processing tasks. We show that our incremental learning method can boost the learning time by up to 20x without compromising the quality of the learned network representation.