Hyper-structure mining of frequent patterns in uncertain data streams.

Hyper-structure mining of frequent patterns in uncertain data streams.
复制标题

DOI:
10.1007/s10115-012-0581-y
复制
发表时间:
2013-10-01
影响因子:
2.7
通讯作者:
Tu, Yi-cheng
Tu, Yi-cheng
中科院分区:
计算机科学4区
文献类型:
--
作者:
HewaNadungodage, Chandima;Xia, Yuni;Lee, Jaehwan John;Tu, Yi-cheng

文献摘要

参考文献

被引文献

相似文献

数据不确定性是许多现实应用中固有的,例如传感器监控系统、基于位置的服务和医疗诊断系统。此外,许多现实世界的应用程序现在能够产生连续的,无限的数据流。近年来,新的方法已经开发出在不确定数据库中发现频繁模式,然而,在不确定数据流中发现频繁模式的工作非常有限。目前不确定流中频繁模式挖掘的解决方案采用基于FP树的方法;然而,最近的研究表明,基于FP树的算法在数据不确定性的存在下表现不佳。在本文中,我们提出了两个超结构为基础的假阳性导向的算法,有效地挖掘频繁项集的不确定数据流。第一个算法,UHS-Stream,被设计为找到所有的频繁项集到当前时刻。第二个算法TFUHS-Stream,被设计为在不确定的数据流中以时间衰落的方式发现频繁项集。实验结果表明,所提出的基于超结构的算法优于现有的基于树的算法在准确性,运行时间和内存使用。
Data uncertainty is inherent in many real-world applications such as sensor monitoring systems, location-based services, and medical diagnostic systems. Moreover, many real-world applications are now capable of producing continuous, unbounded data streams. During the recent years, new methods have been developed to find frequent patterns in uncertain databases; nevertheless, very limited work has been done in discovering frequent patterns in uncertain data streams. The current solutions for frequent pattern mining in uncertain streams take a FP-tree-based approach; however, recent studies have shown that FP-tree-based algorithms do not perform well in the presence of data uncertainty. In this paper, we propose two hyper-structure-based false-positive-oriented algorithms to efficiently mine frequent itemsets from streams of uncertain data. The first algorithm, UHS-Stream, is designed to find all frequent itemsets up to the current moment. The second algorithm, TFUHS-Stream, is designed to find frequent itemsets in an uncertain data stream in a time-fading manner. Experimental results show that the proposed hyper-structure-based algorithms outperform the existing tree-based algorithms in terms of accuracy, runtime, and memory usage.
DOI: 10.1007/s10115-010-0363-3
发表时间: 2012-01-01
影响因子: 2.7
作者:
Salam, Abdus;Khayal, M. Sikandar Hayat
通讯作者: Khayal, M. Sikandar Hayat
DOI: 10.1007/s10115-007-0092-4
发表时间: 2008-07-01
影响因子: 2.7
作者:
Cheng, James;Ke, Yiping;Ng, Wilfred
通讯作者: Ng, Wilfred
挖掘数据流中频繁项的方法:概述
DOI: 10.1007/s10115-009-0267-2
发表时间: 2011-01-01
影响因子: 2.7
作者:
Liu, Hongyan;Lin, Yuan;Han, Jiawei
通讯作者: Han, Jiawei
DOI: 10.1007/s10115-010-0309-9
发表时间: 2011-06-01
影响因子: 2.7
作者:
Yoan Rodriguez-Gonzalez, Ansel;Francisco Martinez-Trinidad, Jose;Ruiz-Shulcloper, Jose
通讯作者: Ruiz-Shulcloper, Jose
DOI: 10.1007/s10115-007-0112-4
发表时间: 2008-10-01
影响因子: 2.7
作者:
Li, Hua-Fu;Shan, Man-Kwan;Lee, Suh-Yin
通讯作者: Lee, Suh-Yin