Bump hunting in high-dimensional data

Bump hunting in high-dimensional data
复制标题

DOI:
10.1023/a:1008894516817
复制
发表时间:
1999-04-01
影响因子:
2.2
通讯作者:
Fisher, NI
Fisher, NI
中科院分区:
数学2区
文献类型:
--
作者:
Friedman, JH;Fisher, NI

文献摘要

被引文献

相似文献

许多数据分析问题可表述为(有噪声的)优化问题。它们明确或隐含地涉及寻找一组(“输入”)变量值的同时组合,这些组合意味着另一个指定(“输出”)变量具有异常大(或小)的值。具体而言,人们寻求输入变量空间的一组子区域,在这些子区域内输出变量的值比其在整个输入域上的平均值大得多(或小得多)。此外,通常希望这些区域能够以一种可解释的形式描述,涉及关于输入值的简单陈述(“规则”)。本文基于“耐心”规则归纳的概念提出了一个旨在实现这一目标的过程。这种耐心策略与大多数规则归纳方法所使用的贪心策略以及一些分区树技术(如CART)所使用的半贪心策略形成对比。文中还介绍了涉及科学和商业数据库的应用。
Many data analytic questions can be formulated as (noisy) optimization problems. They explicitly or implicitly involve finding simultaneous combinations of values for a set of ("input") variables that imply unusually large (or small) values of another designated ("output") variable. Specifically, one seeks a set of subregions of the input variable space within which the value of the output variable is considerably larger (or smaller) than its average value over the entire input domain. In addition it is usually desired that these regions be describable in an interpretable form involving simple statements ("rules") concerning the input values. This paper presents a procedure directed towards this goal based on the notion of "patient" rule induction. This patient strategy is contrasted with the greedy ones used by most rule induction methods, and semi-greedy ones used by some partitioning tree techniques such as CART. Applications involving scientific and commercial data bases are presented.