Discovering Interesting Holes in Data

Discovering Interesting Holes in Data
复制标题

发现数据中有趣的漏洞

DOI:
--
复制
发表时间:
1997
期刊:
International Joint Conference on Artificial Intelligence
影响因子:
--
通讯作者:
W. Hsu
W. Hsu
中科院分区:
--
文献类型:
--
作者:
B. Liu;Liang;W. Hsu

文献摘要

被引文献

相似文献

当前的机器学习和发现技术侧重于发现数据中存在的规则或规律。过去被忽视的研究的一个重要方面是学习或发现数据库中有趣的漏洞。如果我们将数据库中的每个案例视为it维空间中的一个点,那么空洞就是空间中不包含数据点的一个区域。显然,不是每个洞都有趣。有些漏洞是显而易见的,因为已知某些值组合是不可能的。存在一些漏洞是因为数据库中没有足够的案例。然而,在某些情况下,空白区域确实携带着重要的信息。例如,它们可以警告我们一些缺失的值组合,这些组合要么是以前不知道的,要么是意想不到的。了解这些缺失的价值组合可能会带来重大发现。在本文中,我们提出了一种发现数据库漏洞的算法。
Current machine learning and discovery techniques focus on discovering rules or regularities that exist in data. An important aspect of the research that has been ignored in the past is the learning or discovering of interesting holes in the database. If we view each case in the database as a point in a it-dimensional space, then a hole is simply a region in the space that contains no data point. Clearly, not every hole is interesting. Some holes are obvious because it is known that certain value combinations are not possible. Some holes exist because there are insufficient cases in the database. However, in some situations, empty regions do carry important information. For instance, they could warn us about some missing value combinations that are either not known before or are unexpected. Knowing these missing value combinations may lead to significant discoveries. In this paper, we propose an algorithm to discover holes in databases.