PYTHIA -

PYTHIA -
复制标题

皮提亚 -

DOI:
--
复制
发表时间:
2007
期刊:
影响因子:
--
通讯作者:
John R. RiceAbstractKnowledge
John R. RiceAbstractKnowledge
中科院分区:
--
文献类型:
--
作者:
C. Houstis;Vassilis S. Verykios;Ann C. Catlin;Naren Ramakrishnan;John R. RiceAbstractKnowledge

文献摘要

被引文献

相似文献

数据库中的知识发现(KDD)是一个新兴的、不断发展的研究领域,它试图通过自动获取隐藏在现实生活中的操作数据库中存储的大量数据中的知识来解决知识获取瓶颈。这种归纳知识的方法已应用于手动检查不可行的各种领域。在我们的研究中,我们采用并适当修改了这种方法,并将其应用于存储与科学计算应用解决方案相关的性能数据的数据库。有人认为,科学数据库适合自动机器检查,因为存储的信息质量良好,不会丢失值或不一致的数据。本文介绍的系统 PYTHIA-II 通过将强大的 DBMS 无缝集成到知识获取过程中,为知识工程师提供了组织和存储数据的能力。灵活且全面的图形用户界面简化了模式提取相关数据的选择,该界面支持用户与数据挖掘工具集的交互。该系统使用挖掘工具生成的知识结构来构建知识库,该知识库支持科学计算领域的推荐系统的功能。特别是,本文描述了 KDD 过程在涉及椭圆偏微分方程求解软件评估的案例研究中的应用。 1. 简介 数据库中的知识发现(KDD)是一个新兴的跨学科领域,旨在发现大型、现实的、可操作的数据库系统中隐藏的信息。 KDD 方法中吸引该领域大多数研究人员兴趣的阶段是数据挖掘。在此阶段,数据挖掘算法应用于目标数据集,以发现将用于构建底层域模型的模式。这项研究解决的三个最重要的问题是:必须在短时间内处理的大量数据、这些系统中包含的不完整和脏信息以及数据的时变性质。对于第一个问题,研究人员寻求可以通过并行系统有效实现的可扩展的知识发现方法,以及可以以灵活和最佳的方式从永久数据存储中访问数据的先进技术。 “不完整和脏”数据应由挖掘算法本身处理;随时间变化的数据需要对已发现的知识进行增量更新。
Knowledge Discovery in Databases (KDD) is a new and evolving research area which attempts to solve the knowledge acquisition bottleneck by automatically acquiring knowledge hidden in enormous amounts of data stored in real-life, operational databases. This methodology of inducing knowledge has been applied to a variety of domains where manual inspection was not feasible. In our research, we have adopted, appropriately modiied, and applied this methodology to databases storing performance data related to the solution of scientiic computing applications. It has been argued that scientiic databases lend themselves to automatic machine inspection, since the stored information is of good quality-without missing values or inconsistent data. PYTHIA-II, the system presented in this paper, gives a knowledge engineer the capability of organizing and storing data by seemlessly integrating a powerful DBMS into the knowledge aquisition process. The selection of relevant data for pattern extraction is simpliied by a exible and comprehensive graphical user interface which supports user interaction with a collection of data mining tools. The system uses knowledge structures generated by the mining tools to build a knowledge base which supports the functionality of a recommender system for the scientiic computing domain. In particular, this paper describes the application of the KDD process to a case study involving the evaluation of software for the solution of elliptic Partial Diierential Equations. 1. INTRODUCTION Knowledge Discovery in Databases (KDD) is an emerging, interdisciplinary eld that seeks to uncover hidden information in large, real-life, operational database systems. The phase of the KDD methodology that has attracted the interest of a majority of researchers in this area is Data Mining. During this phase, a data mining algorithm is applied to a target set of data to uncover patterns that will be used in building a model of the underlying domain. Three of the most important issues addressed by this research are: the enormous amount of data that must be processed in a short period of time, the incomplete and dirty" information contained in these systems and the time varying nature of the data. With respect to the rst issue, researchers seek scalable knowledge discovery methodologies that can be implemented eeciently by parallel systems, as well as advanced techniques that can access data from permanent data stores in a exible and optimal manner. Incomplete and dirty" data should be handled by the mining algorithms themselves; time varying data calls for incremental updates in the discovered knowledge.