A database perspective on knowledge discovery

A database perspective on knowledge discovery
复制标题

DOI:
10.1145/240455.240472
复制
发表时间:
1996-11-01
影响因子:
22.7
通讯作者:
Mannila, H
Mannila, H
中科院分区:
计算机科学3区
文献类型:
--
作者:
Imielinski, T;Mannila, H

文献摘要

被引文献

相似文献

ACM通讯1996年11月/第39卷,第11号,59详细说明。对于大多数在线分析处理(OLAP)工具来说也是如此。我们称这样的系统为第一代数据库挖掘系统。当前的情况与20世纪60年代早期数据库管理系统的情况非常相似,当时每个应用程序都必须从头开始构建,没有后来由SQL和关系数据库API(应用程序编程接口)提供的专用数据库原语的好处。事实上,今天的数据挖掘技术更适合描述为“文件挖掘”,因为它们假设数据挖掘引擎和数据库之间的松散耦合。在大多数情况下,该接口具有两个简单命令的形式:“读取”和“写入”,就像30年前Cobol程序与大型数据文件交互一样。此外,KDD领域目前缺乏身份:对于KDD是否是一个独立的领域没有达成共识。一些人声称知识发现只是对大型数据集的机器学习,KDD的数据库组件本质上是最大化在大型持久数据集顶部运行的挖掘操作的性能,并涉及昂贵的I/O。当然,缩小归纳学习工具和现有数据库系统之间的差距是一项重要的工作。然而,尽管提高性能是一个重要问题,但它可能不足以引发系统能力的质的变化。此外,许多数据库管理系统(DBMS)的性能增强数据库挖掘的重要性也是可取的。这些特性包括并行查询执行、内存评估以及对采样和聚合查询操作(如“计数"、“平均”和“方差”)的优化支持。这些特性具有普遍的意义,它们对数据库挖掘的重要性只能被视为DBMS基础研究的副作用。此外,现有DBMS的增量改进,以更好地适应KDD应用程序将很可能是不够的,因为目前的DBMS主要针对不同类别的应用程序。一个历史的类比可能是合适的;有人可能会说,性能的改进,单独的I/O操作将永远不会引发DBMS的研究领域在近30年前。查询语言、查询优化和事务处理是过去三十年数据库领域巨大增长背后的驱动思想。更具体地说,正是查询的特殊性质给构建通用查询优化器带来了挑战。如果查询是预定义的,并且数量有限,那么开发高度调优的独立库例程就足够了。我们相信数据库挖掘可以遵循类似的发展道路。
COMMUNICATIONS OF THE ACM November 1996/Vol. 39, No. 11 59 covery feature. The same is true for most online analytical processing (OLAP) tools. We call such systems firstgeneration database mining systems. The current situation is very similar to the situation in database management systems in the early 1960s, when each application had to be built from scratch, without the benefit of dedicated database primitives provided later by SQL and relational database APIs (Application Programming Interface). In fact, today’s techniques of data mining would more appropriately be described as “file mining” since they assume a loose coupling between a data-mining engine and a database. In most cases, this interface has a form of two simple commands:“read from” and “write to”, just as Cobol programs interacted with large data files 30 years ago. Additionally, the KDD field currently suffers from lack of identity: there is no consensus on whether KDD is an area of its own. Some claim knowledge discovery is simply machine learning with large data sets, and that the database component of the KDD is essentially maximizing performance of mining operations running in the top of large persistent data sets and involving expensive I/O. It is, of course, important work to close the gap between the inductive learning tools and the database systems currently available. However, although improving performance is an important issue, it is probably not sufficient to trigger a qualitative change in system capabilities. Moreover, many database management system (DBMS) performance enhancements important for database mining are also desirable in general. Such features include parallel query execution, in-memory evaluation as well as optimized support for sampling and aggregate query operations such as “Count,”“Average,” and “Variance.” These features are of general interest and their importance to database mining can be viewed only as a side effect of basic research in DBMS. Moreover, incremental improvements of existing DBMSs to better suit KDD applications will most likely be insufficient, since the current DBMSs were primarily targeted at the different classes of applications.A historical analogy is perhaps in order; one may argue that performance improvement of I/O operations alone would have never triggered the DBMS research field nearly 30 years ago. Query languages, query optimization, and transaction processing were the driving ideas behind the tremendous growth of database field in the last three decades. To be more specific, it was the ad hoc nature of querying that created a challenge to build general-purpose query optimizers. If queries were predefined and their number was limited it would be sufficient to develop highly tuned, stand-alone, library routines. We believe database mining can follow a similar path of development.