Topology Preserving Data Reduction for Computing Persistent Homology

Topology Preserving Data Reduction for Computing Persistent Homology
复制标题

DOI:
10.1109/bigdata50022.2020.9378216
复制
发表时间:
2020-12
期刊:
2020 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Nicholas O. Malott;Aaron M. Sens;P. Wilsey
Nicholas O. Malott;Aaron M. Sens;P. Wilsey
中科院分区:
其他
文献类型:
--
作者:
Nicholas O. Malott;Aaron M. Sens;P. Wilsey

文献摘要

相似文献

一种新兴的数据分析方法被称为拓扑数据分析(TDA)。TDA基于拓扑学的数学领域,并研究连续变形下的空间属性。用于TDA的关键工具之一被称为持久同源性,它考虑不同空间分辨率下d维点云中点的连通性,以识别空间中的拓扑属性(孔,环和空隙)。持久同源性,然后分类的拓扑特征,通过它们的持久性的空间连接性的范围。不幸的是,计算持久同源性的内存和运行时复杂度是指数级的,当前的工具只能处理$\mathbb{R}^{3}$中的几千个点。幸运的是,数据简化技术的使用使得持久同源性能够应用于更大的点云。减少数据的技术范围从点的随机采样到聚类数据并使用聚类质心作为减少的数据。虽然几种数据简化方法似乎保留了原始点云中存在的大拓扑特征,但没有系统的研究比较不同数据聚类技术在保留持久同源性结果方面的功效。本文探讨了拓扑保持数据约简的问题,并正式描述了何时以及如何拓扑特征可以被错误描述或丢失的数据约简技术。本文还进行了实验评估的数据减少技术和弹性的持久同源性的影响。特别是,通过随机选择的数据减少相比,从不同的数据聚类算法提取的聚类质心。
An emerging method for data analysis is called Topological Data Analysis (TDA). TDA is based in the mathematical field of topology and examines the properties of spaces under continuous deformation. One of the key tools used for TDA is called persistent homology which considers the connectivity of points in a d-dimensional point cloud at different spatial resolutions to identify topological properties (holes, loops, and voids) in the space. Persistent homology then classifies the topological features by their persistence through the range of spatial connectivity. Unfortunately the memory and run-time complexity of computing persistent homology is exponential and current tools can only process a few thousand points in $\mathbb{R}^{3}$. Fortunately, the use of data reduction techniques enables persistent homology to be applied to much larger point clouds. Techniques to reduce the data range from random sampling of points to clustering the data and using the cluster centroids as the reduced data. While several data reduction approaches appear to preserve the large topological features present in the original point cloud, no systematic study comparing the efficacy of different data clustering techniques in preserving the persistent homology results has been performed. This paper explores the question of topology preserving data reductions and describes formally when and how topological features can be mischaracterized or lost by data reduction techniques. The paper also performs an experimental assessment of data reduction techniques and resilient effects on the persistent homology. In particular, data reduction by random selection is compared to cluster centroids extracted from different data clustering algorithms.