Learning metrics for persistence-based summaries and applications for graph classification

Learning metrics for persistence-based summaries and applications for graph classification
复制标题

DOI:
--
复制
发表时间:
2019-04
期刊:
ArXiv
影响因子:
--
通讯作者:
Qi Zhao;Yusu Wang
Qi Zhao;Yusu Wang
中科院分区:
其他
文献类型:
--
作者:
Qi Zhao;Yusu Wang

文献摘要

被引文献

相似文献

最近,一种新的特征表示和数据分析方法开始受到关注,该方法基于一种称为持久同调的拓扑工具(及其相应的持久化图摘要)。为了便于机器学习工具的后续使用,已经开发了一系列方法来将持久性图映射到向量表示,并且在这些方法中,不同持久性特征的重要性(权重)通常是预先设定的。然而,在实践中,权重函数的选择通常取决于所考虑的特定数据类型的性质,因此,从标记数据中学习最佳权重函数(以及持久性图的度量)是非常可取的。我们研究了这个问题,并开发了一个新的加权内核,称为WKPI,用于持久性摘要,以及一个优化框架,以学习持久性摘要的良好度量。我们的核问题和优化问题都有很好的性质。我们进一步将学习到的核应用于具有挑战性的图分类任务,并表明我们基于wkpi的分类框架在一系列基准数据集上获得了与以前的一系列图分类框架相似或(有时显着)更好的结果。
Recently a new feature representation and data analysis methodology based on a topological tool called persistent homology (and its corresponding persistence diagram summary) has started to attract momentum. A series of methods have been developed to map a persistence diagram to a vector representation so as to facilitate the downstream use of machine learning tools, and in these approaches, the importance (weight) of different persistence features are often preset. However often in practice, the choice of the weight function should depend on the nature of the specific type of data one considers, and it is thus highly desirable to learn a best weight function (and thus metric for persistence diagrams) from labelled data. We study this problem and develop a new weighted kernel, called WKPI, for persistence summaries, as well as an optimization framework to learn a good metric for persistence summaries. Both our kernel and optimization problem have nice properties. We further apply the learned kernel to the challenging task of graph classification, and show that our WKPI-based classification framework obtains similar or (sometimes significantly) better results than the best results from a range of previous graph classification frameworks on a collection of benchmark datasets.