AutoGDA: Automated Graph Data Augmentation for Node Classification

AutoGDA: Automated Graph Data Augmentation for Node Classification
复制标题

DOI:
--
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Tong Zhao;Xianfeng Tang;Danqing Zhang;Haoming Jiang;Nikhil S. Rao;Yiwei Song;Pallav Agrawal;Karthik Subbian;Bing Yin;Meng Jiang
Tong Zhao;Xianfeng Tang;Danqing Zhang;Haoming Jiang;Nikhil S. Rao;Yiwei Song;Pallav Agrawal;Karthik Subbian;Bing Yin;Meng Jiang
中科院分区:
其他
文献类型:
--
作者:
Tong Zhao;Xianfeng Tang;Danqing Zhang;Haoming Jiang;Nikhil S. Rao;Yiwei Song;Pallav Agrawal;Karthik Subbian;Bing Yin;Meng Jiang

文献摘要

相似文献

图数据扩充被用来提高图机器学习的泛化能力。然而,通过仅在整个图上应用固定的扩增操作,现有方法忽略了自然存在于图中的社区的独特特征。例如,不同的社区可以有不同的度分布和同质比。忽略整个图上的统一艾德增强策略的这种差异可能导致图数据增强方法的次优性能。在本文中,我们研究了一个新的问题,自动图数据增强节点分类从本地化的角度来看,社区。我们将其表述为一个双层优化问题:为每个社区找到一组增强策略,最大限度地提高图神经网络在节点分类上的性能。由于两层优化难以直接求解,而社区自定义增强策略的搜索空间巨大,本文提出了一个强化学习框架AutoGDA,该框架可以顺序学习每个社区的局部最优增强策略。我们提出的方法在公共节点分类基准以及真实的行业电子商务网络上的表现优于已建立和流行的基线,准确率高达+12.5%。
Graph data augmentation has been used to improve generalizability of graph machine learning. However, by only applying fixed augmentation operations on entire graphs, existing methods overlook the unique characteristics of communities which naturally exist in the graphs. For example, different communities can have various degree distributions and homophily ratios. Ignoring such discrepancy with unified augmentation strategies on the entire graph could lead to sub-optimal performance for graph data augmentation methods. In this paper, we study a novel problem of automated graph data augmentation for node classification from the localized perspective of communities. We formulate it as a bilevel optimization problem: finding a set of augmentation strategies for each community, which maximizes the performance of graph neural networks on node classification. As the bilevel optimization is hard to solve directly and the search space for community-customized augmentations strategy is huge, we propose a reinforcement learning framework AutoGDA that learns the local-optimal augmentation strategy for each community sequentially. Our proposed approach outperforms established and popular baselines on public node classification benchmarks as well as real industry e-commerce networks by up to +12.5% accuracy.