Topics in Predictive and Descriptive Data Mining
Topics in Predictive and Descriptive Data Mining
批准号:
0204029
负责人:
Jerome Friedman
金额:
$42.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2002
资助国家:
美国
项目状态:
已结题
起止时间:
2002-07-01 至 2008-06-30
中文摘要
摘要PI:曾傑瑞弗里德曼DMS-0204029本提案寻求对预测性和描述性数据挖掘(DM)研究的支持。决策树方法是最流行的预测性数据挖掘工具。在这项资助下的研究将探索克服其最严重限制的方法:在存在具有非常大量的值(因素)的分类(因素)预测变量时严重过度匹配。聚类分析通常被用作描述性数据挖掘的一种工具。在大多数数据挖掘应用中,每次观测都要测量大量的变量。通常,聚类,如果存在的话,只发生在所有测量变量的未知子集中(通常是很小的)。此外,单独的集群可以表示可变子集上的分组(可能重叠)。目标是确定聚类组以及每个组优先聚类的特定变量子集。传统的聚类算法不能很好地适应这一任务。这项拨款下的研究将探索解决这一问题的新方法,特别是在有非常大量可测量变量的情况下。数据挖掘用于发现数据中的模式和关系,重点是非常大的数据库。它在商业、工业、科学、医学以及最近的国土安全方面产生了重大影响。数据挖掘活动分为两种类型:预测性和描述性。预测性数据挖掘涉及使用一个系统过去的观测数据来建立该系统的数学模型。该模型被用来预测系统未来的一些未知属性(属性或变量),给出未来将会知道的其他属性。描述性数据挖掘试图构建紧凑的、可解释的数据摘要,以便理解模式和关系,而不是侧重于对特定属性的预测。这项研究将探索新的方法,以增加描述性和预测性数据挖掘在传统上薄弱的问题上的能力。
英文摘要
AbstractPI: Jerry FriedmanDMS-0204029This proposal seeks support for research in predictive and descriptive data mining (DM). Decision tree methods are the most popular predictive DM tools. Research under this grant will investigate ways to overcome their most serious limitation: severe over fitting in the presence of categorical (factorial) predictor variables with very large numbers of values (factors). Cluster analysis is often used as a tool in descriptive DM. In most DM applications a large number of variables are measured on each observation. Usually clustering, if it exists, occurs only within (often small) unknown subsets of all the measured variables. Moreover, individual clusters may represent groupings on (possibly overlapping) variable subsets. The goal is to identify the clustered groups as well as the particular variable subsets on which each one preferentially clusters. Traditional clustering algorithms are not well suited for this task. Research under this grant will investigate new approaches for solving this problem, especially in situations where there are a very large number of measured variables.Data mining is used to discover patterns and relationships in data, with an emphasis on very large data bases. It has had a major impact in business, industry, science, medicine, and most recently homeland security. Data mining activities divide into two types: predictive and descriptive. Predictive DM involves using past observational data from a system to build a mathematical model of that system. The model is used to predict some future unknown property (attribute or variable) of the system, given other properties that will be known in the future. Descriptive DM seeks to construct compact, interpretable summaries of the data in order to understand patterns and relationships, without focusing on the prediction of particular attributes. This research will investigate new methodologies for increasing the power of both descriptive and predictive DM in problems for which they have been traditionally weak.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
New Directions in Predictive Learning for Classification
-
批准号:9704431
-
项目类别:Continuing Grant
-
资助金额:$39.87万
-
财政年份:1997
-
负责人:Jerome Friedman
-
依托单位:
Mathematical Sciences: Adaptive Spatial Regression and Classification
-
批准号:9403804
-
项目类别:Continuing Grant
-
资助金额:$11.0万
-
财政年份:1994
-
负责人:Jerome Friedman
-
依托单位:
海外基金