课题基金 / 基金详情

A clustering framework for the process of knowledge discovery in databases

A clustering framework for the process of knowledge discovery in databases
数据库中知识发现过程的聚类框架
批准号:
250960-2006
负责人:
Ester, Martin
金额:
$2.23万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2008
资助国家:
加拿大
项目状态:
已结题
起止时间:
2008-01-01 至 2009-12-31

项目摘要

项目成果

Ester, Martin的其他基金

相似基金

相关文献

中文摘要
翻译
数据库中的知识发现被定义为从大型数据库中提取有效的、新颖的、可理解的和潜在有用的模式的过程。知识发现过程涉及几个步骤,特别是聚焦、预处理、数据挖掘和评估,这些步骤通常必须反复进行才能取得令人满意的结果。到目前为止,大多数KDD研究都集中在数据挖掘步骤上,为聚类、分类和关联规则挖掘等任务开发了高效的算法。不幸的是,没有太多的研究涉及其他步骤和整个过程,这严重限制了现有数据挖掘方法的有效性。在这个项目中,我们希望探索在集群的背景下对整个KDD过程的支持,这是最重要的数据挖掘任务之一。缺乏对所有KDD步骤的支持对于集群来说尤其有问题,因为它是无监督的、探索性的,并且因为大多数集群算法不生成显式模式,而只是将集群简单地作为对象集返回。这一拟议项目的目标是开发一个支持所有知识发现步骤的集群化框架。大多数现有的聚类算法只利用待聚类对象的属性,但在许多新兴应用中,目标表和相关表属性之间的关系在表示感兴趣的对象方面扮演着重要的角色。例如,在市场细分中,不仅购买偏好,而且客户之间的社交网络也与集群相关。作为另一个例子,当对基因表达数据进行聚类时,必须考虑相关蛋白质的属性及其进一步的关系。我们计划与领域专家密切合作,在基因表达数据分析、流式细胞仪数据分析以及社区识别和市场细分的应用程序中对我们的集群框架进行评估。
英文摘要
Knowledge Discovery in Databases (KDD) has been defined as the process of extracting valid, novel, understandable and potentially useful patterns from large databases. The KDD process involves several steps, in particular focusing, pre-processing, data mining and evaluation, that normally have to be iterated to achieve satisfactory results. So far, most KDD research has focused on the data mining step, developing efficient algorithms for tasks such as clustering, classification and association rule mining. Unfortunately, not much research has addressed the other steps and the process as a whole, which has seriously limited the usefulness of existing data mining methods. In this project, we want to explore support for the entire KDD process in the context of clustering, one of the most important data mining tasks. The lack of support for all KDD steps is especially problematic for clustering due to its unsupervised, exploratory nature and because most clustering algorithms do not generate explicit patterns, but return clusters simply as sets of objects. The objective of this proposed project is to develop a framework for clustering supporting all KDD steps. Most existing clustering algorithms exploit only attributes of the objects to be clustered, but in many emerging applications relationships among the target table and attributes from related tables play an important role in representing the objects of interest. In market segmentation, e.g., not only the purchasing preferences but also the social network among the customers is relevant for clustering. When clustering gene expression data, as another example, attributes of the related proteins and their further relationships must be considered. We plan to evaluate our clustering framework in close collaboration with domain experts in the applications of analysis of gene expression data, analysis of flow cytometry data as well as community identification and market segmentation.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Data Mining in Heterogeneous Information Networks with Attributes
  • 批准号:
    RGPIN-2017-04072
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $3.06万
  • 财政年份:
    2022
  • 负责人:
    Ester, Martin
  • 依托单位:
Data Mining in Heterogeneous Information Networks with Attributes
  • 批准号:
    RGPIN-2017-04072
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $3.06万
  • 财政年份:
    2021
  • 负责人:
    Ester, Martin
  • 依托单位:
Data Mining in Heterogeneous Information Networks with Attributes
  • 批准号:
    RGPIN-2017-04072
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $3.06万
  • 财政年份:
    2020
  • 负责人:
    Ester, Martin
  • 依托单位:
Data Mining in Heterogeneous Information Networks with Attributes
  • 批准号:
    RGPIN-2017-04072
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $3.06万
  • 财政年份:
    2019
  • 负责人:
    Ester, Martin
  • 依托单位:
海外基金