Capacity and Bias of Learned Geometric Embeddings for Directed Graphs

Capacity and Bias of Learned Geometric Embeddings for Directed Graphs
复制标题

DOI:
--
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
Michael Boratko;Dongxu Zhang;Nicholas Monath;L. Vilnis;K. Clarkson;A. McCallum
Michael Boratko;Dongxu Zhang;Nicholas Monath;L. Vilnis;K. Clarkson;A. McCallum
中科院分区:
其他
文献类型:
--
作者:
Michael Boratko;Dongxu Zhang;Nicholas Monath;L. Vilnis;K. Clarkson;A. McCallum

文献摘要

被引文献

相似文献

各种各样的机器学习任务,如知识库完成、本体比对和多标签分类fi正离子,都可以使fi受益于将图形或分类的可区分表示纳入学习。虽然欧氏空间中的矢量在理论上可以表示任何图形,但最近的工作表明,诸如复数、双曲线、顺序或盒嵌入等备选方案具有更适合于模拟真实世界的图形的几何性质。然而,在实验上,这些收益只在较低的维度上才能看到,而性能优势fi在较高的维度上会减少。在这项工作中,我们引入了一种新的盒嵌入变体,它使用一个学习的平滑参数来获得比低维向量模型更好的表示能力,同时也避免了其他高维几何模型常见的性能饱和。此外,我们还给出了证明盒嵌入可以表示任意DAG的理论结果。我们在几类合成和真实世界有向图上对向量、双曲和基于区域的几何表示进行了严格的经验评估。对这些结果的分析揭示了不同图族、图特征、模型大小和嵌入几何之间的相关性,为各种可微图表示的归纳偏差提供了有用的见解。
A wide variety of machine learning tasks such as knowledge base completion, ontology alignment, and multi-label classification can benefit from incorporating into learning differentiable representations of graphs or taxonomies. While vectors in Euclidean space can theoretically represent any graph, much recent work shows that alternatives such as complex, hyperbolic, order, or box embeddings have geometric properties better suited to modeling real-world graphs. Experimentally these gains are seen only in lower dimensions, however, with performance benefits diminishing in higher dimensions. In this work, we introduce a novel variant of box embeddings that uses a learned smoothing parameter to achieve better representational capacity than vector models in low dimensions, while also avoiding performance saturation common to other geometric models in high dimensions. Further, we present theoretical results that prove box embeddings can represent any DAG. We perform rigorous empirical evaluations of vector, hyperbolic, and region-based geometric representations on several families of synthetic and real-world directed graphs. Analysis of these results exposes correlations between different families of graphs, graph characteristics, model size, and embedding geometries, providing useful insights into inductive biases of various differentiable graph representations.