课题基金 / 基金详情

High Performance Rough Sets Data Analysis in Data Mining

High Performance Rough Sets Data Analysis in Data Mining
数据挖掘中的高性能粗糙集数据分析
批准号:
0514750
负责人:
Yi Pan
金额:
$13.75万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2005
资助国家:
美国
项目状态:
已结题
起止时间:
2005-07-15 至 2010-06-30

项目摘要

项目成果

Yi Pan的其他基金

相似基金

相关文献

中文摘要
翻译
数据挖掘(又名数据库中的知识发现,KDD)是从庞大的数据集中提取以前未知的、潜在有用的信息或模式的过程。KDD通常是一个多阶段的过程,涉及数据准备、数据预处理、特征选择、规则归纳、知识评估和部署等多个步骤。许多新的数据挖掘和学习算法虽然很活跃,但在相当多的临时和模糊的概念下被开发出来。在大多数情况下,这些算法是不同研究人员的个人创造,没有太多共同的方法和基本框架。换句话说,数据挖掘的大部分工作都集中在算法开发上,而忽视了关于数据、数据之间的关系以及隐藏在数据中的隐含信息或数据冗余的质量等基本理论问题的研究。因此,要充分理解和评估各个阶段是如何相互影响的,以及每个阶段对整个知识发现过程的影响是不容易的。为了进一步发展和突破数据挖掘和学习算法,有必要对其基础进行深入研究。该研究的中心目标是开发一个统一的基于粗糙集的数据挖掘框架,以探索数据挖掘和学习算法的各种基本问题。它的目的是在数据挖掘方法、技术和应用的背景下展示粗糙集方法的分析能力。它将提供一个统一的框架,帮助更好地理解整个知识发现过程。智力优势:粗糙集理论特别适合于对不精确或不完整的数据进行推理,并发现数据中的关系。粗糙集理论的简单性和数学清晰度使其对理论家和面向应用的研究者都很有吸引力。粗糙集理论的主要优点是它不需要关于数据的任何初始或附加信息,如统计学中的概率、Dempster-Shafer理论中的基本概率赋值或模糊集理论中的隶属度。粗糙集理论为知识发现奠定了良好的基础,可以应用于知识发现过程的不同阶段。特别是,粗糙集理论的形式化技术在属性函数或部分函数依赖及其发现、分析和表征、特征选择、特征提取、数据约简、决策规则生成和模式提取(模板、关联规则)等方面产生了许多新颖而有前途的突破性方法和算法,这些都是知识发现过程的基本问题。粗糙集理论代表了一种新的创新方法,可以导致新的学习算法的开发,从而创造数据挖掘技术的新用途和突破。广泛的影响:拟议的协作项目本质上是跨学科的。它将综合数据挖掘、粗糙集理论和高性能计算中经常不同的工作。PIS强大的多学科研究协作经验将使拟议的研究对粗糙集、数据挖掘和高性能计算社区产生广泛的认识和影响。它将在一个统一的框架内设计和开发一系列新的数据挖掘算法和方法,包括数据约简、规则归纳和分类集成,以更好地理解整个KDD过程。这些算法和方法将极大地扩展数据挖掘技术和粗糙集理论的应用范围,并将使人们更好地理解设计高效和创新的数据挖掘和学习算法和方法所涉及的问题。建议的研究将与教学活动紧密结合,研究成果将发展为本科生和研究生课程和研究项目。这一方法的一部分包括开发新的跨学科课程,将计算机科学和数学结合在一起,以了解数据挖掘和粗糙集理论基础的原理和方法。这种集成将有助于培训学生在粗糙集理论所涉及的问题,设计和实现新的数据挖掘方法和算法,高性能计算。学生的积极参与将使他们有机会接触到数据挖掘领域的最新研究。
英文摘要
Data mining (aka Knowledge Discovery in Databases, KDD) is a procedure to extract previously unknown and potentially useful information or pattern from huge data sets. KDD is usually a multiphase process involving numerous steps such as data preparation, data preprocessing, feature selection, rule induction, knowledge evaluation and deployment etc. Many novel data mining and learning algorithms have been developed, though vigorously, under rather add hoc and vague concepts. These algorithms, in most cases, are individual creations of different researchers, without much common methodological and fundamental framework. In other words, great majority of work in data mining is focused on algorithm development while neglecting the studies of fundamental theoretical issues concerning data, inter-data relationships, and quality of the implicit information hidden in the data or data redundancies. Thus, it is not easy to fully understand and evaluate how individual phase influences each other and the impact of each phase on the whole knowledge discovery process. For further development and breakthroughs in data mining and learning algorithms, a deep examination of its foundation is necessary. The central goal of the proposed research is to develop a unified rough set based data mining framework to explore various fundamental issues of data mining and learning algorithms. It aims to present the analytical capabilities of the methodology of rough sets in the context of data mining methodologies, techniques and applications. It will provide a unified framework to help better understand the whole KDD process.Intellectual merit: Rough set theory is particularly suited to reasoning about imprecise or incomplete data and discovering relationships in the data. The simplicity and mathematical clarity of rough set theory makes it attractive for both theoreticians and application-oriented researchers. The main advantage of rough set theory is that it does not require any preliminary or additional information about the data, such as probability in statistics, basic probability assignment in Dempster-Shafer theory or the value of membership in fuzzy set theory. Rough set theory constitutes a sound basis for KDD and can be used in different phases of the KDD process. In particular, the formal techniques of rough set theory lead to many novel and promising breakthrough methods and algorithms for attribute functional, orpartial functional dependencies, their discovery, analysis, and characterization, feature election, feature extraction, data reduction, decision rule generation, and pattern extraction (templates, association rules) etc., which are the fundamental issues of the KDD process. Rough set theory represents a new innovative approach and can lead to the development of new learning algorithms to create novel uses and breakthroughs of data mining techniques.Broader impacts: The proposed collaborative project is interdisciplinary in nature. It will synthesize often-disparate work in data mining, rough set theory and high performance computing. The PIs' strong multidisciplinary research collaboration experience will lead to widespread awareness and impact of the proposed research to rough set, data mining and high performance computing community. It will design and develop a wide-range of novel data mining algorithms and methods including data reduction, rule induction and classification ensemble in one unified framework to better understand the whole KDDprocess. These algorithms and methods will significantly extend the application scope of data mining techniques and rough set theory and will result in the improved understanding of issues involved in designing efficient and innovative data mining and learning algorithms and methods. The proposed research will integrate tightly with teaching activities, the research results will be developed into undergraduate and graduate courses and research projects. Part of this approach includes the development of new cross-disciplinary courses that bring together computer science and mathematics for the understanding of principle and methods of theoretical foundations of data mining and rough set theory. The integration will help with training students in the issues involved in the rough set theory, design and implementation of novel data mining methods and algorithms, high performance computing. The active participation of students will allow for significant exposure to the latest research in datamining.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Real World Relevant Security Labware for Mobile Threat Analysis and Protection Experience
Capacity Building: Collaborative Research: Integrated Learning Environment for Cyber Security of Smart Grid
Travel Awards for The 2011 IEEE International Conference on Bioinformatics & Biomedicine
(NECO) Collaborative Research: Reliability Modeling for Large-Scale Networking System (LSNS), and Self-Improvement in LSNS
国内基金
海外基金
基于Rough Path理论的分布依赖随机微分方程的平均化原理研究
Rough随机波动率模型的金融应用及算法研究
  • 批准号:
    12071373
  • 项目类别:
    面上项目
  • 资助金额:
    52.0万元
  • 批准年份:
    2020
  • 负责人:
    马敬堂
  • 依托单位:
带跳的 rough path 理论及其应用
  • 批准号:
    11901104
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    27.0万元
  • 批准年份:
    2019
  • 负责人:
    张会林
  • 依托单位:
基于Rough集的坚硬顶板条件下煤与瓦斯突出预警机制研究
  • 批准号:
    51874121
  • 项目类别:
    面上项目
  • 资助金额:
    60.0万元
  • 批准年份:
    2018
  • 负责人:
    杨玉中
  • 依托单位: