Visual management of large scale data mining projects.

Visual management of large scale data mining projects.
复制标题

大规模数据挖掘项目的可视化管理。

DOI:
10.1142/9789814447331_0026
复制
发表时间:
2000
影响因子:
--
通讯作者:
Hunter,L
Hunter,L
中科院分区:
--
文献类型:
--
作者:
Shah,I;Hunter,L

文献摘要

被引文献

相似文献

本文描述了一个统一的框架,用于可视化数百个机器学习实验的准备和结果。这些实验旨在提高从序列预测酶功能的准确性,在许多情况下都是成功的。我们的系统提供了图形用户界面,用于定义和探索训练数据集和各种代表性备选方案,用于检查由各种类型的学习算法得出的假设,用于可视化全局结果,以及用于详细检查特定训练集(功能)和示例(蛋白质)的结果。可视化工具在大量序列数据和归纳知识中起到导航辅助作用。它们为理解我们成功和失败的意义和潜在的生物学解释提供了重要的帮助。利用这些可视化,可以有效地识别模块序列表示和归纳算法的弱点,从而提出更好的学习策略。我们的数据挖掘可视化工具包的开发背景是从蛋白质序列数据中准确预测酶功能的问题。以前的工作9表明,仅基于序列相似性,大约6%的酶蛋白序列可能被赋予错误的功能。为了验证这样的假设,即使用机器学习技术和模块化领域表示进行更详细的序列分析可以解决许多这样的失败,我们设计了一系列超过250个实验,使用信息论决策树归纳和朴素贝叶斯学习对有问题的酶功能类的局部序列领域表示进行了研究。在其中一半以上的情况下,我们的方法能够完美地区分相似序列的各种可能功能10。我们在此应用程序上开发并测试了我们的可视化技术。
This paper describes a unified framework for visualizing the preparations for, and results of, hundreds of machine learning experiments. These experiments were designed to improve the accuracy of enzyme functional predictions from sequence, and in many cases were successful. Our system provides graphical user interfaces for defining and exploring training datasets and various representational alternatives, for inspecting the hypotheses induced by various types of learning algorithms, for visualizing the global results, and for inspecting in detail results for specific training sets (functions) and examples (proteins). The visualization tools serve as a navigational aid through a large amount of sequence data and induced knowledge. They provided significant help in understanding both the significance and the underlying biological explanations of our successes and failures. Using these visualizations it was possible to efficiently identify weaknesses of the modular sequence representations and induction algorithms which suggest better learning strategies. The context in which our data mining visualization toolkit was developed was the problem of accurately predicting enzyme function from protein sequence data. Previous work9demonstrated that approximately 6% of enzyme protein sequences are likely to be assigned incorrect functions on the basis of sequence similarity alone. In order to test the hypothesis that more detailed sequence analysis using machine learning techniques and modular domain representations could address many of these failures, we designed a series of more than 250 experiments using information-theoretic decision tree induction and naive Bayesian learning on local sequence domain representations of problematic enzyme function classes. In more than half of these cases, our methods were able to perfectly discriminate among various possible functions of similar sequences10. We developed and tested our visualization techniques on this application.