Two Sides of the Same Coin: Heterophily and Oversmoothing in Graph Convolutional Neural Networks

Two Sides of the Same Coin: Heterophily and Oversmoothing in Graph Convolutional Neural Networks
复制标题

DOI:
10.1109/icdm54844.2022.00169
复制
发表时间:
2021-02
期刊:
2022 IEEE International Conference on Data Mining (ICDM)
影响因子:
--
通讯作者:
Yujun Yan;Milad Hashemi;Kevin Swersky;Yaoqing Yang;Danai Koutra
Yujun Yan;Milad Hashemi;Kevin Swersky;Yaoqing Yang;Danai Koutra
中科院分区:
其他
文献类型:
--
作者:
Yujun Yan;Milad Hashemi;Kevin Swersky;Yaoqing Yang;Danai Koutra

文献摘要

被引文献

相似文献

在节点分类任务中,图卷积神经网络(GCN)在不同的图数据上表现出了优于传统方法的性能。然而,众所周知,GCN 的性能会随着层数的增加而降低(过度平滑问题),并且最近的研究还表明,GCN 在异亲图中可能表现更差,其中相邻节点往往属于不同的类(异亲问题)。这两个问题通常被视为不相关,因此是独立研究的,通常是从谱角度在图过滤器级别进行研究。我们是第一个以统一的视角共同解释节点级别的过平滑和异质性问题的人。具体来说,我们通过两个定量指标来分析节点:节点的相对程度(与其邻居相比)和节点级异质性。我们的理论表明,这两个分析指标的相互作用定义了节点行为的三种情况,它们共同解释了过度平滑和异质性问题,并可以预测 GCN 的性能。基于我们理论的见解,我们从理论上和经验上展示了两种策略的有效性:基于结构的边缘校正,从结构属性(即度)学习校正的边缘权重,以及基于特征的边缘校正,从节点特征学习带符号的边缘权重。与其他能够很好地处理异质性或过度平滑的方法相比,我们表明,我们的模型 GGCN 结合了这两种策略,在这两个问题上都表现良好。我们在 [1] 中提供了本文的较长版本,并在 https://github.com/YujunYan/Heterophily_and_oversmoothing 上提供了代码。
In node classification tasks, graph convolutional neural networks (GCNs) have demonstrated competitive performance over traditional methods on diverse graph data. However, it is known that the performance of GCNs degrades with increasing number of layers (oversmoothing problem) and recent studies have also shown that GCNs may perform worse in heterophilous graphs, where neighboring nodes tend to belong to different classes (heterophily problem). These two problems are usually viewed as unrelated, and thus are studied independently, often at the graph filter level from a spectral perspective.We are the first to take a unified perspective to jointly explain the oversmoothing and heterophily problems at the node level. Specifically, we profile the nodes via two quantitative metrics: the relative degree of a node (compared to its neighbors) and the node-level heterophily. Our theory shows that the interplay of these two profiling metrics defines three cases of node behaviors, which explain the oversmoothing and heterophily problems jointly and can predict the performance of GCNs. Based on insights from our theory, we show theoretically and empirically the effectiveness of two strategies: structure-based edge correction, which learns corrected edge weights from structural properties (i.e., degrees), and feature-based edge correction, which learns signed edge weights from node features. Compared to other approaches, which tend to handle well either heterophily or oversmoothing, we show that our model, GGCN, which incorporates the two strategies performs well in both problems. We provide a longer version of this paper in [1] and codes on https://github.com/YujunYan/Heterophily_and_oversmoothing.