An Information Theoretic Approach to Detection of Minority Subsets in Database

An Information Theoretic Approach to Detection of Minority Subsets in Database
复制标题

DOI:
10.1109/icdm.2006.19
复制
发表时间:
2006-12
期刊:
Sixth International Conference on Data Mining (ICDM'06)
影响因子:
--
通讯作者:
S. Ando;Einoshin Suzuki
S. Ando;Einoshin Suzuki
中科院分区:
其他
文献类型:
--
作者:
S. Ando;Einoshin Suzuki

文献摘要

相似文献

大规模数据库中罕见和异常事件的检测已经成为知识发现和信息检索领域的一项重要实践。许多数据库包含大量噪声或无关数据,其分布往往与包含有用知识的特殊数据的子集重叠。本文讨论的问题是找到少数数据的一小部分,其分布与数据库的大多数数据重叠,但与大多数数据库的数据例外或不一致。在这种情况下,传统的基于距离或基于密度的孤立点检测方法由于依赖于大多数的结构或关键参数的前提而无效。我们将该任务形式化为对少数子集模型的估计,该模型提供了对子集的简单描述,但又保持了与多数子集的发散。利用率失真理论的信息论框架,将这种估计形式化为最小化问题。我们进一步引入多数条件,得到了一个目标函数,该函数分解了少数的性质和对多数结构的依赖。该方法在人工数据方面较传统方法有了很大的改进,在文档检索问题上也取得了良好的效果。
Detection of rare and exceptional occurrences in large- scale databases have become an important practice in the field of knowledge discovery and information retrieval. Many databases include large amount of noise or irrelevant data, whose distribution often overlaps with the subsets of exceptional data containing useful knowledge. This paper addresses the problem of finding a small subset of "minority" data whose distribution overlaps with, but are exceptional to or inconsistent with that of the majority of the database. In such a case, conventional distance-based or density-based approaches in Outlier Detection are ineffective due to their dependence on the structure of the majority or the prerequisite of critical parameters. We formalize the task as an estimation of a model of the minority subset which provides a simple description of the subset and yet maintains divergence from that of the majority. This estimation is formalized as a minimization problem using an information theoretic framework of Rate Distortion theory. We further introduce conditions of the majority to derive an objective function which factorizes the property of the minority and dependence to the structure of the majority. The proposed method shows improvements from conventional approaches in artificial data and a promising result in document retrieval problem.