Theoretical and empirical analysis of ReliefF and RReliefF

Theoretical and empirical analysis of ReliefF and RReliefF
复制标题

DOI:
10.1023/a:1025667309714
复制
发表时间:
2003-10-01
期刊:
影响因子:
7.5
通讯作者:
Kononenko, I
Kononenko, I
中科院分区:
计算机科学3区
文献类型:
--
作者:
Robnik-Sikonja, M;Kononenko, I

文献摘要

被引文献

相似文献

救济算法是通用的、成功的属性估计器。它们能够检测属性之间的条件依赖关系,并为回归和分类中的属性估计提供统一的视图。此外,他们的质量估计也有一种自然的解读。虽然它们通常被认为是在学习模型之前的前置步骤中应用的特征子集选择方法,但它们实际上已经被成功地用于各种设置,例如,在决策或回归树学习的构建阶段选择分裂或指导建构性归纳,作为属性加权方法,以及在归纳逻辑编程中。在本文中,我们从理论和经验两个方面调查和讨论了它们的工作原理和原因,它们的理论和实用特性,它们的参数,它们检测到的哪种依赖关系,它们如何扩展到大量的样本和特征,如何为它们采样数据,它们对噪声的稳健性如何,无关和冗余的属性如何影响它们的输出,以及不同的度量标准如何影响它们。
Relief algorithms are general and successful attribute estimators. They are able to detect conditional dependencies between attributes and provide a unified view on the attribute estimation in regression and classification. In addition, their quality estimates have a natural interpretation. While they have commonly been viewed as feature subset selection methods that are applied in prepossessing step before a model is learned, they have actually been used successfully in a variety of settings, e. g., to select splits or to guide constructive induction in the building phase of decision or regression tree learning, as the attribute weighting method and also in the inductive logic programming.A broad spectrum of successful uses calls for especially careful investigation of various features Relief algorithms have. In this paper we theoretically and empirically investigate and discuss how and why they work, their theoretical and practical properties, their parameters, what kind of dependencies they detect, how do they scale up to large number of examples and features, how to sample data for them, how robust are they regarding the noise, how irrelevant and redundant attributes influence their output and how different metrics influences them.