A Multiscale Environment for Learning by Diffusion

A Multiscale Environment for Learning by Diffusion
复制标题

DOI:
10.1016/j.acha.2021.11.004
复制
发表时间:
2021-01
期刊:
ArXiv
影响因子:
--
通讯作者:
James M. Murphy;Sam L. Polk
James M. Murphy;Sam L. Polk
中科院分区:
其他
文献类型:
--
作者:
James M. Murphy;Sam L. Polk

文献摘要

相似文献

聚类算法将数据集划分为多组相似的点。集群问题非常普遍,同一数据集的不同分区可以被认为是正确和有用的。为了充分理解这些数据,必须从各种不同的尺度来考虑,从粗略到精细。我们介绍了多尺度扩散学习环境(MELD)数据模型,它是一类通过数据集上的非线性扩散来参数化的聚类族。结果表明,MELD数据模型准确地捕捉了数据中潜在的多尺度结构,有利于数据的分析。为了有效地学习在实际数据集中观察到的多尺度结构,我们引入了无监督非线性扩散多尺度学习(M-LUND)聚类算法,该算法是从时间尺度上的扩散过程派生出来的。我们为算法的性能提供了理论保证,并证明了算法的计算效率。最后,我们展示了M-Lund聚类算法在一系列合成和真实数据集中检测潜在结构。
Clustering algorithms partition a dataset into groups of similar points. The clustering problem is very general, and different partitions of the same dataset could be considered correct and useful. To fully understand such data, it must be considered at a variety of scales, ranging from coarse to fine. We introduce the Multiscale Environment for Learning by Diffusion (MELD) data model, which is a family of clusterings parameterized by nonlinear diffusion on the dataset. We show that the MELD data model precisely captures latent multiscale structure in data and facilitates its analysis. To efficiently learn the multiscale structure observed in many real datasets, we introduce the Multiscale Learning by Unsupervised Nonlinear Diffusion (M-LUND) clustering algorithm, which is derived from a diffusion process at a range of temporal scales. We provide theoretical guarantees for the algorithm's performance and establish its computational efficiency. Finally, we show that the M-LUND clustering algorithm detects the latent structure in a range of synthetic and real datasets.