A database perspective on knowledge discovery
A database perspective on knowledge discovery
复制标题
DOI:
10.1145/240455.240472
复制
发表时间:
1996-11-01
影响因子:
22.7
通讯作者:
Mannila, H
中科院分区:
文献类型:
--
作者:
Imielinski, T;Mannila, H
COMMUNICATIONS OF THE ACM November 1996/Vol. 39, No. 11 59 covery feature. The same is true for most online analytical processing (OLAP) tools. We call such systems firstgeneration database mining systems. The current situation is very similar to the situation in database management systems in the early 1960s, when each application had to be built from scratch, without the benefit of dedicated database primitives provided later by SQL and relational database APIs (Application Programming Interface). In fact, today’s techniques of data mining would more appropriately be described as “file mining” since they assume a loose coupling between a data-mining engine and a database. In most cases, this interface has a form of two simple commands:“read from” and “write to”, just as Cobol programs interacted with large data files 30 years ago. Additionally, the KDD field currently suffers from lack of identity: there is no consensus on whether KDD is an area of its own. Some claim knowledge discovery is simply machine learning with large data sets, and that the database component of the KDD is essentially maximizing performance of mining operations running in the top of large persistent data sets and involving expensive I/O. It is, of course, important work to close the gap between the inductive learning tools and the database systems currently available. However, although improving performance is an important issue, it is probably not sufficient to trigger a qualitative change in system capabilities. Moreover, many database management system (DBMS) performance enhancements important for database mining are also desirable in general. Such features include parallel query execution, in-memory evaluation as well as optimized support for sampling and aggregate query operations such as “Count,”“Average,” and “Variance.” These features are of general interest and their importance to database mining can be viewed only as a side effect of basic research in DBMS. Moreover, incremental improvements of existing DBMSs to better suit KDD applications will most likely be insufficient, since the current DBMSs were primarily targeted at the different classes of applications.A historical analogy is perhaps in order; one may argue that performance improvement of I/O operations alone would have never triggered the DBMS research field nearly 30 years ago. Query languages, query optimization, and transaction processing were the driving ideas behind the tremendous growth of database field in the last three decades. To be more specific, it was the ad hoc nature of querying that created a challenge to build general-purpose query optimizers. If queries were predefined and their number was limited it would be sufficient to develop highly tuned, stand-alone, library routines. We believe database mining can follow a similar path of development.