Ricci Curvature-Based Graph Sparsification for Continual Graph Representation Learning

Ricci Curvature-Based Graph Sparsification for Continual Graph Representation Learning
复制标题

用于连续图表示学习的基于 Ricci 曲率的图稀疏化

DOI:
10.1109/tnnls.2023.3303454
复制
发表时间:
2024
影响因子:
10.4
通讯作者:
Tao, Dacheng
Tao, Dacheng
中科院分区:
计算机科学1区
文献类型:
--
作者:
Zhang, Xikun;Song, Dongjin;Tao, Dacheng

文献摘要

相似文献

内存重放存储以前任务的历史数据子集,以便在学习新任务时重放,为欧几里得数据上的各种持续学习应用程序展示了最先进的性能。虽然拓扑信息在表征图数据方面起着至关重要的作用,但现有的基于内存重放的图学习技术仅存储单个节点以进行重放,并且不考虑其相关的边缘信息。为此,基于图神经网络(GNN)中的消息传递机制,我们提出了基于 Ricci 曲率的图稀疏化技术来执行连续的图表示学习。具体来说,我们首先开发子图情景存储器(SEM),以计算子图的形式存储拓扑信息。接下来,我们稀疏子图,使其仅包含信息最丰富的结构(节点和边)。信息量是用里奇曲率来评估的,这是一种理论上合理的度量,用于估计邻居对表示目标节点的贡献。通过这种方式,我们可以减少计算子图 fromto 的内存消耗,并使 GNN 能够充分利用信息最丰富的拓扑信息进行内存重放。此外,为了保证在大图上的适用性,我们还为稀疏化过程中的里奇曲率提供了理论上合理的代理,这可以极大地方便计算。最后,我们的实证研究表明,SEM 在四个不同的公共数据集上显着优于最先进的方法。与主要关注任务增量学习(task-IL)设置的现有方法不同,SEM 在具有挑战性的类增量学习(class-IL)设置中也取得了成功,其中模型需要区分所有没有任务指标的学习类,甚至达到与联合训练相当的性能,这是持续学习的性能上限。
Memory replay, which stores a subset of historical data from previous tasks to replay while learning new tasks, exhibits state-of-the-art performance for various continual learning applications on the Euclidean data. While topological information plays a critical role in characterizing graph data, existing memory replay-based graph learning techniques only store individual nodes for replay and do not consider their associated edge information. To this end, based on the message-passing mechanism in graph neural networks (GNNs), we present the Ricci curvature-based graph sparsification technique to perform continual graph representation learning. Specifically, we first develop the subgraph episodic memory (SEM) to store the topological information in the form of computation subgraphs. Next, we sparsify the subgraphs such that they only contain the most informative structures (nodes and edges). The informativeness is evaluated with the Ricci curvature, a theoretically justified metric to estimate the contribution of neighbors to represent a target node. In this way, we can reduce the memory consumption of a computation subgraph fromtoand enable GNNs to fully utilize the most informative topological information for memory replay. Besides, to ensure the applicability on large graphs, we also provide the theoretically justified surrogate for the Ricci curvature in the sparsification process, which can greatly facilitate the computation. Finally, our empirical studies show that SEM outperforms state-of-the-art approaches significantly on four different public datasets. Unlike existing methods, which mainly focus on task incremental learning (task-IL) setting, SEM also succeeds in the challenging class incremental learning (class-IL) setting in which the model is required to distinguish all learned classes without task indicators and even achieves comparable performance to joint training, which is the performance upper bound for continual learning.