A roadmap for the computation of persistent homology

A roadmap for the computation of persistent homology
复制标题

DOI:
10.1140/epjds/s13688-017-0109-5
复制
发表时间:
2017-08-09
期刊:
影响因子:
3.6
通讯作者:
Harrington, Heather A.
Harrington, Heather A.
中科院分区:
计算机科学3区
文献类型:
--
作者:
Otter, Nina;Porter, Mason A.;Harrington, Heather A.

文献摘要

被引文献

相似文献

持久同源性(PH)是拓扑数据分析(TDA)中使用的一种方法,用于研究跨多个尺度持久存在的数据的定性特征。它对输入数据的扰动是鲁棒的,与维度和坐标无关,并且提供了输入的定性特征的紧凑表示。PH的计算是一个开放的领域,有许多重要和迷人的挑战。PH计算领域正在快速发展,新的算法和软件实现正在快速更新和发布。我们的文章的目的是(1)向广泛的计算科学家介绍PH的理论和计算方法,(2)提供PH计算的最先进实现的基准。我们对PH进行了友好的介绍,导航PH计算的管道,着眼于应用,并使用一系列的合成和真实世界的数据集,以评估目前可用的开源实现的PH值的计算。基于我们的基准,我们指出哪些算法和实现是最适合不同类型的数据集。在附带的教程中,我们提供了PH计算的指导方针。我们公开了我们为教程编写的所有脚本,并提供了基准测试中使用的数据集的处理版本。
Persistent homology (PH) is a method used in topological data analysis (TDA) to study qualitative features of data that persist across multiple scales. It is robust to perturbations of input data, independent of dimensions and coordinates, and provides a compact representation of the qualitative features of the input. The computation of PH is an open area with numerous important and fascinating challenges. The field of PH computation is evolving rapidly, and new algorithms and software implementations are being updated and released at a rapid pace. The purposes of our article are to (1) introduce theory and computational methods for PH to a broad range of computational scientists and (2) provide benchmarks of state-of-the-art implementations for the computation of PH. We give a friendly introduction to PH, navigate the pipeline for the computation of PH with an eye towards applications, and use a range of synthetic and real-world data sets to evaluate currently available open-source implementations for the computation of PH. Based on our benchmarking, we indicate which algorithms and implementations are best suited to different types of data sets. In an accompanying tutorial, we provide guidelines for the computation of PH. We make publicly available all scripts that we wrote for the tutorial, and we make available the processed version of the data sets used in the benchmarking.