Fast outlier data mining algorithm based on cell in large datasets

Fast outlier data mining algorithm based on cell in large datasets
复制标题

DOI:
--
复制
发表时间:
2010
期刊:
Journal of Chongqing University of Posts and Telecommunications
影响因子:
--
通讯作者:
Gou Guang-le
Gou Guang-le
中科院分区:
其他
文献类型:
--
作者:
Gou Guang-le

文献摘要

被引文献

相似文献

本文提出了一种快速的基于单元格的大型数据集离群点检测算法(简称formabcld)。该算法采用聚类技术对数据进行预处理,根据数据的值将数据放入相应的单元格中,并用单元格维树对非空单元格进行索引。过滤掉大部分位于高密度单元格中且与离群值没有密切关系的数据,避免了大量无用的计算。实验表明,该方法能够快速、准确地从大型数据集中挖掘离群数据,提高了离群数据的检测速度。
The paper proposed a fast cell-based algorithm for outlier detection in large datasets(short for FOMABCLD).The algorithm applied cluster technique to preprocesse data,and placed data into the appropriate cells based on their values and indexed the non-empty cells with Cell Dimension-Tree.A majority part of data located in high density cells and had no nearness relationship with outliers is filtered,which avoided large useless computations.The experiment show that FOMABCLD can mine outlier data from large datasets fast and accurately,and the speed of detecting outliers is increased.