Information-Theoretic Measures for Knowledge Discovery and Data Mining

Information-Theoretic Measures for Knowledge Discovery and Data Mining
复制标题

DOI:
10.1007/978-3-540-36212-8_6
复制
发表时间:
2003
期刊:
--
影响因子:
--
通讯作者:
Yiyu Yao
Yiyu Yao
中科院分区:
其他
文献类型:
--
作者:
Yiyu Yao

文献摘要

被引文献

相似文献

数据库可以被视为统计总体,属性可以被视为从其域中获取值的统计变量。人们可以对数据库进行统计和信息论分析。根据属性值,可以将数据库划分为更小的群体。如果某个属性对数据库进行了分区,使得可以观察到以前未知的规律和模式,则该属性被认为是重要的。人们已经提出并应用了许多信息论方法来量化各个领域中属性的重要性以及属性之间的关系。在知识发现和数据挖掘(KDD)的背景下,我们对属性重要性和属性关联的信息论度量进行了批判性的回顾和分析,重点是它们的解释和联系。
A database may be considered as a statistical population, and an attribute as a statistical variable taking values from its domain. One can carry out statistical and information-theoretic analysis on a database. Based on the attribute values, a database can be partitioned into smaller populations. An attribute is deemed important if it partitions the database such that previously unknown regularities and patterns are observable. Many information-theoretic measures have been proposed and applied to quantify the importance of attributes and relationships between attributes in various fields. In the context of knowledge discovery and data mining (KDD), we present a critical review and analysis of information-theoretic measures of attribute importance and attribute association, with emphasis on their interpretations and connections.