Classification and Novel Class Detection in Concept-Drifting Data Streams under Time Constraints

Classification and Novel Class Detection in Concept-Drifting Data Streams under Time Constraints
复制标题

DOI:
10.1201/9781315119458-13
复制
发表时间:
2011-06
影响因子:
8.9
通讯作者:
M. Masud;Jing Gao;L. Khan;Jiawei Han;B. Thuraisingham
M. Masud;Jing Gao;L. Khan;Jiawei Han;B. Thuraisingham
中科院分区:
计算机科学2区
文献类型:
--
作者:
M. Masud;Jing Gao;L. Khan;Jiawei Han;B. Thuraisingham

文献摘要

被引文献

相似文献

大多数现有的数据流分类技术忽略了流数据的一个重要方面:一个新的类的到来。我们解决这个问题,并提出了一种数据流分类技术,集成了一种新的类检测机制到传统的分类器,使新的类的真实标签到达之前,新的类的自动检测。新的类检测问题变得更具挑战性的概念漂移的存在下,当底层数据分布演变的流。为了确定一个实例是否属于一个新的类,分类模型有时需要等待更多的测试实例来发现这些实例之间的相似性。施加最大允许等待时间Tc作为对测试实例进行分类的时间约束。此外,大多数现有的流分类方法假设数据点的真实标签可以在数据点被分类之后立即被访问。实际上,在获得数据点的真实标记时涉及时间延迟T1,因为手动标记是耗时的。我们展示了如何在这些约束下做出快速正确的分类决策,并将其应用于真实的基准数据。与最先进的流分类技术的比较证明了我们的方法的优越性。
Most existing data stream classification techniques ignore one important aspect of stream data: arrival of a novel class. We address this issue and propose a data stream classification technique that integrates a novel class detection mechanism into traditional classifiers, enabling automatic detection of novel classes before the true labels of the novel class instances arrive. Novel class detection problem becomes more challenging in the presence of concept-drift, when the underlying data distributions evolve in streams. In order to determine whether an instance belongs to a novel class, the classification model sometimes needs to wait for more test instances to discover similarities among those instances. A maximum allowable wait time Tc is imposed as a time constraint to classify a test instance. Furthermore, most existing stream classification approaches assume that the true label of a data point can be accessed immediately after the data point is classified. In reality, a time delay Tl is involved in obtaining the true label of a data point since manual labeling is time consuming. We show how to make fast and correct classification decisions under these constraints and apply them to real benchmark data. Comparison with state-of-the-art stream classification techniques prove the superiority of our approach.