A Framework for Clustering Uncertain Data Streams

A Framework for Clustering Uncertain Data Streams
复制标题

DOI:
10.1109/icde.2008.4497423
复制
发表时间:
2008-04
期刊:
2008 IEEE 24th International Conference on Data Engineering
影响因子:
--
通讯作者:
C. Aggarwal;Philip S. Yu
C. Aggarwal;Philip S. Yu
中科院分区:
其他
文献类型:
--
作者:
C. Aggarwal;Philip S. Yu

文献摘要

被引文献

相似文献

近年来,由于大量硬件应用程序近似测量数据,不确定数据管理应用程序变得越来越重要。例如,由于数据检索、传输和电源故障的不准确性,传感器的读数通常会产生相当大的噪声。在许多情况下,底层数据流的估计误差是可用的。该信息对于挖掘过程非常有用,因为它可以用来提高基础结果的质量。在本文中,我们将提出一种对不确定数据流进行聚类的方法。我们使用一个非常通用的不确定性模型,其中我们假设只有少数不确定性的统计度量可用。我们将证明,在挖掘过程中使用即使是适度的不确定性信息也足以大大提高潜在结果的质量。我们证明我们的方法比纯粹的确定性方法(例如 CluStream 方法)更有效。我们将在各种真实和合成数据集上测试该方法,并说明该方法在有效性和效率方面的优势。
In recent years, uncertain data management applications have grown in importance because of the large number of hardware applications which measure data approximately. For example, sensors are typically expected to have considerable noise in their readings because of inaccuracies in data retrieval, transmission, and power failures. In many cases, the estimated error of the underlying data stream is available. This information is very useful for the mining process, since it can be used in order to improve the quality of the underlying results. In this paper we will propose a method for clustering uncertain data streams. We use a very general model of the uncertainty in which we assume that only a few statistical measures of the uncertainty are available. We will show that the use of even modest uncertainty information during the mining process is sufficient to greatly improve the quality of the underlying results. We show that our approach is more effective than a purely deterministic method such as the CluStream approach. We will test the approach on a variety of real and synthetic data sets and illustrate the advantages of the method in terms of effectiveness and efficiency.