Data Clustering Using Topological Features

Data Clustering Using Topological Features
复制标题

使用拓扑特征的数据聚类

DOI:
--
复制
发表时间:
2014
期刊:
Brazilian Conference on Intelligent Systems
影响因子:
--
通讯作者:
R. Mello
R. Mello
中科院分区:
--
文献类型:
--
作者:
Cássio M. M. Pereira;R. Mello

文献摘要

被引文献

相似文献

聚类是最常用的数据挖掘技术之一,而计算拓扑学是一个非常新的领域,将抽象数学与具体的计算技术联系起来。在本文中,我们探讨的假设,拓扑相似的集群可能表明有意义的关系。我们的方法有一个有效的实现计算最小生成树的基础上,以获得每个集群的拓扑信息。然后,我们计算离散性和不连通性指数,用于表征每个集群,从而允许检索等价类。我们表明,对于一个真实世界的高维网络入侵数据集,我们的方法检索到的拓扑相似的集群确实对应于有意义的等价类存在于数据集中。
Clustering is one of the most used data mining techniques, while computational topology is a very recent field bridging abstract mathematics with concrete computational techniques. In this paper, we explore the hypothesis that topologically-similar clusters may indicate meaningful relationships. Our approach has an efficient implementation based on computing Minimum Spanning Trees to obtain topological information of each cluster. We then compute a discreteness and a disconnectedness index, used to characterize each cluster, thus allowing the retrieval of equivalence classes. We show that for a real-world high-dimensional network intrusion data set, the topologically-similar clusters retrieved by our approach do indeed correspond to meaningful equivalence classes present in the data set.