A statistical interpretation of spectral embedding: The generalised random dot product graph

A statistical interpretation of spectral embedding: The generalised random dot product graph
复制标题

DOI:
10.1111/rssb.12509
复制
发表时间:
2017-09
期刊:
Journal of the Royal Statistical Society: Series B (Statistical Methodology)
影响因子:
--
通讯作者:
Patrick Rubin-Delanchy;C. Priebe;M. Tang;Joshua Cape
Patrick Rubin-Delanchy;C. Priebe;M. Tang;Joshua Cape
中科院分区:
其他
文献类型:
--
作者:
Patrick Rubin-Delanchy;C. Priebe;M. Tang;Joshua Cape

文献摘要

被引文献

相似文献

谱嵌入是一种可以用来获得图的节点的向量表示的过程。本文提出了一个概括的潜在位置网络模型称为随机点积图,允许解释这些向量表示为潜在的位置估计。一般化是需要模型heterophilic连接(例如“异性吸引”),以科普更普遍的负特征值。我们发现,无论是使用的邻接或归一化拉普拉斯矩阵,谱嵌入产生一致的潜在位置估计渐近高斯误差(可识别性)。标准和混合隶属度随机块模型是特殊情况,其中潜在位置仅取K个不同的向量值,代表社区,或分别存在于具有这些顶点的(K − 1)-单纯形中。在随机块模型下,我们的理论建议使用高斯混合模型(而不是K均值)进行谱聚类,并且在混合隶属度下,拟合包围单纯形的最小体积,现有的建议以前只支持非负定假设。在网络安全示例中展示了链接预测(在随机点积图上)的经验改进,以及揭示更丰富的潜在结构(比标准或混合成员随机块模型下的假设)的潜力。
Spectral embedding is a procedure which can be used to obtain vector representations of the nodes of a graph. This paper proposes a generalisation of the latent position network model known as the random dot product graph, to allow interpretation of those vector representations as latent position estimates. The generalisation is needed to model heterophilic connectivity (e.g. ‘opposites attract’) and to cope with negative eigenvalues more generally. We show that, whether the adjacency or normalised Laplacian matrix is used, spectral embedding produces uniformly consistent latent position estimates with asymptotically Gaussian error (up to identifiability). The standard and mixed membership stochastic block models are special cases in which the latent positions take only K distinct vector values, representing communities, or live in the (K − 1)‐simplex with those vertices respectively. Under the stochastic block model, our theory suggests spectral clustering using a Gaussian mixture model (rather than K‐means) and, under mixed membership, fitting the minimum volume enclosing simplex, existing recommendations previously only supported under non‐negative‐definite assumptions. Empirical improvements in link prediction (over the random dot product graph), and the potential to uncover richer latent structure (than posited under the standard or mixed membership stochastic block models) are demonstrated in a cyber‐security example.