Incremental density-based ensemble clustering over evolving data streams

Incremental density-based ensemble clustering over evolving data streams
复制标题

在不断变化的数据流上进行基于增量密度的集成聚类

DOI:
10.1016/j.neucom.2016.01.009
复制
发表时间:
2016-05
期刊:
影响因子:
6
通讯作者:
Kamen Ivanov
Kamen Ivanov
中科院分区:
计算机科学2区
文献类型:
--
作者:
Imran Khan;Joshua Zhexue Huang;Kamen Ivanov

文献摘要

参考文献

被引文献

相似文献

智能电表技术的最新进展使实时收集客户用电量信息成为可能。测量值是连续产生的,在某些情况下,例如在工业智能计量中,数据交换率是高度波动的。这种具有大量缺失值和稀疏值的智能电表流数据的存储、查询和挖掘是非常具有计算挑战性的任务。为了解决这些问题,我们提出了一种新的方法,称为基于密度的增量集成聚类(IDEStream),用于基于电力消耗数据的各种工厂的增量分割。它利用伽马混合模型来抑制数据流中稀疏数据单元的影响,这些数据流在一个时间窗口内依次到达,然后从该窗口的处理数据生成聚类。IDEStream使用一种独特的增量集成方法,对后续时间窗口的聚类进行增量聚合。在中国广东省制造工厂的智能电表收集的数据流上的实验结果表明,该算法优于几种最先进的数据流聚类算法。获得的细分可以找到许多应用,一个典型的例子是以灵活的方式定义客户费率。
The recent advances in smart meter technology have enabled for collecting information about customer power consumption in real time. The measurements are generated continuously and in some cases, e.g. in the industrial smart metering the data exchange rates are highly-fluctuating. The storage, querying, and mining of such smart meter streaming data with a large number of missing and sparse values are highly computationally challenging tasks. To address such matters, we propose a new method called incremental density-based ensemble clustering (IDEStream) for incremental segmentation of various kinds of factories based on their electricity consumption data. It exploits a gamma mixture model to suppress the influence of sparse data units in the data streams that sequentially arrive within a time window and then generates a clustering from the processed data of that window. IDEStream uses a unique incremental ensemble approach to incrementally aggregate the clusterings of subsequent time windows. Experimental results on data streams collected by smart meters from manufacturing factories in Guangdong province of China have shown that the proposed algorithm outperforms several state-of-the-art data stream clustering algorithms. The obtained segmentation can find numerous applications, an exemplar one being to define customer rates in a flexible way.
DOI: 10.1002/9781118445112.stat08170
发表时间: 2019-08
期刊: Wiley StatsRef: Statistics Reference Online
影响因子: --
作者:
A. Acharya;Joydeep Ghosh
通讯作者: A. Acharya;Joydeep Ghosh
DOI: 10.1002/0471722227
发表时间: 2003-07
期刊: --
影响因子: --
作者:
N. Balakrishnan;V. Nevzorov
通讯作者: N. Balakrishnan;V. Nevzorov
DOI: 10.1145/276304.276312
发表时间: 1998-06
期刊: --
影响因子: --
作者:
S. Guha;R. Rastogi;Kyuseok Shim
通讯作者: S. Guha;R. Rastogi;Kyuseok Shim
DOI: 10.1016/0378-4754(87)90094-2
发表时间: 1986
期刊: --
影响因子: --
作者:
坂元 慶行;石黒 真木夫;北川 源四郎
通讯作者: 坂元 慶行;石黒 真木夫;北川 源四郎
DOI: 10.1198/jasa.2004.s341
发表时间: 2004-06
影响因子: 3.7
作者:
C. Anderson‐Cook
通讯作者: C. Anderson‐Cook