Correlation-Based Refinement of Rules with Numerical Attributes

Correlation-Based Refinement of Rules with Numerical Attributes
复制标题

具有数值属性的基于相关性的规则细化

DOI:
--
复制
发表时间:
2014
期刊:
The Florida AI Research Society
影响因子:
--
通讯作者:
Johanna Völker
Johanna Völker
中科院分区:
--
文献类型:
--
作者:
André Melo;M. Theobald;Johanna Völker

文献摘要

被引文献

相似文献

学习规则是提取有用信息的常用方法 来自知识或数据库。很多这样的数据集 包含数字属性。然而,像 ILP 这样的方法 或关联规则挖掘针对具有分类值的数据进行优化,并且考虑数值属性是昂贵的。在本文中,我们提出了自上而下的扩展 ILP 算法(例如 FOIL)可实现高效发现 来自数字和分类数据的规则 属性。我们的方法包括预处理阶段 用于计算数值属性和分类属性之间的相关性,以及 ILP 细化步骤的扩展,这使我们能够检测有趣的候选规则并建议使用相关属性组合进行细化。我们报告了美国人口普查数据、Freebase 和 DBpedia 的实验,并表明我们的方法有助于有效地发现具有数值间隔的规则。
Learning rules is a common way of extracting useful information from knowledge or data bases. Many of such data sets contain numerical attributes. However, approaches like ILP or association rule mining are optimized for data with categorical values, and considering numerical attributes is expensive. In this paper, we present an extension to top-down ILP algorithms such as FOIL, which enables an efficient discovery of rules from data with both numerical and categorical attributes. Our approach comprises a preprocessing phase for computing the correlations between numerical and categorical attributes, as well as an extension to the ILP refinement step, which enables us to detect interesting candidate rules and to suggest refinements with relevant attribute combinations. We report on experiments with U.S. Census data, Freebase and DBpedia, and show that our approach helps to efficiently discover rules with numerical intervals.