Data-Centric Learning from Unlabeled Graphs with Diffusion Model

Data-Centric Learning from Unlabeled Graphs with Diffusion Model
复制标题

DOI:
10.48550/arxiv.2303.10108
复制
发表时间:
2023-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Gang Liu;Eric Inae;Tong Zhao;Jiaxin Xu;Te Luo;Meng Jiang
Gang Liu;Eric Inae;Tong Zhao;Jiaxin Xu;Te Luo;Meng Jiang
中科院分区:
其他
文献类型:
--
作者:
Gang Liu;Eric Inae;Tong Zhao;Jiaxin Xu;Te Luo;Meng Jiang

文献摘要

相似文献

图属性预测任务是重要且众多的。虽然每个任务提供了一个小尺寸的标记的例子,未标记的图形已经从各种来源和大规模收集。传统的方法是在自监督任务上使用未标记图训练模型,然后在预测任务上微调模型。然而,自我监督的任务知识无法与预测所需的内容对齐,有时甚至发生冲突。在本文中,我们提出了提取的知识背后的大型未标记的图形作为一组特定的有用的数据点,以增强每个属性预测模型。我们使用扩散模型来充分利用未标记的图,并设计了两个新的目标来指导模型的去噪过程,每个任务的标记数据生成特定于任务的图示例及其标签。实验表明,我们以数据为中心的方法在15个任务上的性能明显优于15个现有的各种方法。与自监督学习不同,未标记数据带来的性能改进是可见的,因为生成的标记示例。
Graph property prediction tasks are important and numerous. While each task offers a small size of labeled examples, unlabeled graphs have been collected from various sources and at a large scale. A conventional approach is training a model with the unlabeled graphs on self-supervised tasks and then fine-tuning the model on the prediction tasks. However, the self-supervised task knowledge could not be aligned or sometimes conflicted with what the predictions needed. In this paper, we propose to extract the knowledge underlying the large set of unlabeled graphs as a specific set of useful data points to augment each property prediction model. We use a diffusion model to fully utilize the unlabeled graphs and design two new objectives to guide the model's denoising process with each task's labeled data to generate task-specific graph examples and their labels. Experiments demonstrate that our data-centric approach performs significantly better than fifteen existing various methods on fifteen tasks. The performance improvement brought by unlabeled data is visible as the generated labeled examples unlike the self-supervised learning.