A Clustering-based Framework for Classifying Data Streams

A Clustering-based Framework for Classifying Data Streams
复制标题

DOI:
10.24963/ijcai.2021/448
复制
发表时间:
2021-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Xuyang Yan;A. Homaifar;M. Sarkar;Abenezer Girma;E. Tunstel
Xuyang Yan;A. Homaifar;M. Sarkar;Abenezer Girma;E. Tunstel
中科院分区:
其他
文献类型:
--
作者:
Xuyang Yan;A. Homaifar;M. Sarkar;Abenezer Girma;E. Tunstel

文献摘要

相似文献

数据流的非平稳特性对传统的机器学习技术提出了强烈的挑战。虽然已经提出了一些解决方案来扩展传统的机器学习技术来处理数据流,但这些方法要么需要初始标签集,要么依赖于专门的设计参数。类之间的重叠和数据流的标记构成了对数据流进行分类的其他主要挑战。在本文中,我们提出了一个基于聚类的数据流分类框架来处理非平稳数据流,而不使用初始标签集。采用基于密度的流聚类方法捕获具有动态阈值的新概念,并引入有效的主动标签查询策略从数据流中不断学习新概念。探讨了每个簇的子簇结构,以处理类之间的重叠。实验结果和定量比较研究表明,该方法在统计上优于现有方法或具有可比性。
The non-stationary nature of data streams strongly challenges traditional machine learning techniques. Although some solutions have been proposed to extend traditional machine learning techniques for handling data streams, these approaches either require an initial label set or rely on specialized design parameters. The overlap among classes and the labeling of data streams constitute other major challenges for classifying data streams. In this paper, we proposed a clustering-based data stream classification framework to handle non-stationary data streams without utilizing an initial label set. A density-based stream clustering procedure is used to capture novel concepts with a dynamic threshold and an effective active label querying strategy is introduced to continuously learn the new concepts from the data streams. The sub-cluster structure of each cluster is explored to handle the overlap among classes. Experimental results and quantitative comparison studies reveal that the proposed method provides statistically better or comparable performance than the existing methods.