Cluster-based Data Reduction for Persistent Homology

Cluster-based Data Reduction for Persistent Homology
复制标题

DOI:
10.1109/bigdata.2018.8622440
复制
发表时间:
2018-12
期刊:
2018 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Anindya Moitra;Nicholas O. Malott;P. Wilsey
Anindya Moitra;Nicholas O. Malott;P. Wilsey
中科院分区:
其他
文献类型:
--
作者:
Anindya Moitra;Nicholas O. Malott;P. Wilsey

文献摘要

被引文献

相似文献

持久同调用于计算空间在不同空间分辨率下的拓扑特征。它是应用于数据分析问题的计算拓扑学的主要工具之一。尽管多次尝试降低其复杂性,但持续同源性在时间和空间上仍然昂贵。这些限制使得该方法可以应用的最大数据集具有千分之三数量级的点。本文探讨了一种技术,旨在减少数据点的数量,同时保持显着的拓扑特征的数据。所提出的技术,使持久的同源性计算的原始输入数据的简化版本,而不影响输出的重要组成部分。由于持续同源性的运行时间是指数的数据点的数量,所提出的数据减少方法有利于计算在一小部分的时间所需的原始数据。此外,数据简化方法可以与简化持久同源性计算的任何现有技术相结合。通过创建类似数据点的小组(称为纳米集群),然后用其集群中心替换每个纳米集群内的点来执行数据缩减。减少的数据的持久性同源性不同于原始数据的纳米团簇的半径所限定的量。理论分析的实验结果表明,持久的同源性被保存所提出的数据减少技术的支持。
Persistent homology is used for computing topological features of a space at different spatial resolutions. It is one of the main tools from computational topology that is applied to the problems of data analysis. Despite several attempts to reduce its complexity, persistent homology remains expensive in both time and space. These limits are such that the largest data sets to which the method can be applied have the number of points of the order of thousands in ℝ3. This paper explores a technique intended to reduce the number of data points while preserving the salient topological features of the data. The proposed technique enables the computation of persistent homology on a reduced version of the original input data without affecting significant components of the output. Since the run time of persistent homology is exponential in the number of data points, the proposed data reduction method facilitates the computation in a fraction of the time required for the original data. Moreover, the data reduction method can be combined with any existing technique that simplifies the computation of persistent homology. The data reduction is performed by creating small groups of similar data points, called nano-clusters, and then replacing the points within each nano-cluster with its cluster center. The persistence homology of the reduced data differs from that of the original data by an amount bounded by the radius of the nano-clusters. The theoretical analysis is backed by experimental results showing that persistent homology is preserved by the proposed data reduction technique.