Semi-Supervised Graph Imbalanced Regression

Semi-Supervised Graph Imbalanced Regression
复制标题

DOI:
10.1145/3580305.3599497
复制
发表时间:
2023-05
期刊:
Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining
影响因子:
--
通讯作者:
Gang Liu;Tong Zhao;Eric Inae;Te Luo;Meng Jiang
Gang Liu;Tong Zhao;Eric Inae;Te Luo;Meng Jiang
中科院分区:
其他
文献类型:
--
作者:
Gang Liu;Tong Zhao;Eric Inae;Te Luo;Meng Jiang

文献摘要

被引文献

相似文献

在回归任务中,当某些连续标签值的观测数据难以收集时,在标注数据中很容易出现数据不平衡的情况。当涉及分子和聚合物性质预测时,标注的图数据集通常很小,因为对其进行标注需要昂贵的设备和大量的工作。为了解决图回归任务中罕见标签值样本不足的问题,我们提出了一个半监督框架,通过自我训练逐步平衡训练数据并减少模型偏差。训练数据的平衡是通过以下方式实现的:(1)使用一种新的回归置信度度量为代表性不足的标签伪标记更多的图;(2)在使用伪标签平衡数据后,为剩余的罕见标签在潜在空间中扩充图示例。前者是从无标注数据中识别出标签能被可靠预测的优质示例,并从不平衡的标注数据中按照反向分布抽取其中的一个子集。后者与前者协作,使用一种新的基于标签的混合算法来实现完美平衡。我们在图数据集上的七个回归任务中进行了实验。结果表明,所提出的框架显著降低了预测图性质的误差,尤其是在代表性不足的标签区域。
Data imbalance is easily found in annotated data when the observations of certain continuous label values are difficult to collect for regression tasks. When they come to molecule and polymer property predictions, the annotated graph datasets are often small because labeling them requires expensive equipment and effort. To address the lack of examples of rare label values in graph regression tasks, we propose a semi-supervised framework to progressively balance training data and reduce model bias via self-training. The training data balance is achieved by (1) pseudo-labeling more graphs for under-represented labels with a novel regression confidence measurement and (2) augmenting graph examples in latent space for remaining rare labels after data balancing with pseudo-labels. The former is to identify quality examples from unlabeled data whose labels are confidently predicted and sample a subset of them with a reverse distribution from the imbalanced annotated data. The latter collaborates with the former to target a perfect balance using a novel label-anchored mixup algorithm. We perform experiments in seven regression tasks on graph datasets. Results demonstrate that the proposed framework significantly reduces the error of predicted graph properties, especially in under-represented label areas.