Squashing flat files flatter

Squashing flat files flatter
复制标题

将平面文件压扁

DOI:
--
复制
发表时间:
1999
期刊:
Knowledge Discovery and Data Mining
影响因子:
--
通讯作者:
D. Pregibon
D. Pregibon
中科院分区:
--
文献类型:
--
作者:
W. DuMouchel;C. Volinsky;T. Johnson;Corinna Cortes;D. Pregibon

文献摘要

被引文献

相似文献

数据挖掘区别于“经典”机器学习(ML)和统计建模(SM)的一个特征是规模。社会似乎同意这一点,但进展到这一点是有限的。我们提出了一种方法,解决规模在一个新的方式,有可能彻底改变该领域。虽然该方法最直接适用于平面(逐列)数据集,但我们认为它可以适用于其他表示。我们解决这个问题的方法不是扩大单个ML和SM方法。相反,我们更喜欢通过缩小数据集来利用现有方法的整个集合。我们称这种方法为挤压。我们的方法明显优于随机抽样和理论论证表明它如何以及为什么工作得很好。Squashing由三个模块化步骤组成:分组、momentizing和生成(GMG)。这三个步骤描述了压缩管道,其中原始数据(非常大的数据集)被分成相互排斥的组(或箱);在每个组中计算一系列低阶矩;最后将这些矩传递给生成伪数据的例程,该伪数据准确地再现了矩。GMG压缩流水线的结果是具有与原始数据相同的结构的压缩数据集,其中为每个伪数据点添加了反映原始数据到初始组中的分布的权重。任何接受权重的ML或SM方法都可以用于分析加权的伪数据。通过构建,所得分析将模拟原始数据集的相应分析。挤压应该吸引许多子学科,
A feature of data mining that distinguishes it from “classical” machine learning (ML) and statistical modeling (SM) is scale. The community seems to agree on this yet progress to this point has been limited. We present a methodology that addresses scale in a novel fashion that has the potential for revolutionizing the field. While the methodology applies most directly to flat (row by column) data sets we believe that it can be adapted to other representations. Our approach to the problem is not to scale up individual ML and SM methods. Rather we prefer to leverage the entire collection of existing methods by scaling down the data set. We call the method squashing. Our method demonstrably outperforms random sampling and a theoretical argument suggests how and why it works well. Squashing consists of three modular steps: grouping, momentizing, and generating (GMG). These three steps describe the squashing pipeline whereby the original (very large data set) is sectioned off into mutually exclusive groups (or bins); within each group a series of low-order moments are computed; and finally these moments are passed off to a routine that generates pseudo data that accurately reproduce the moments. The result of the GMG squashing pipeline is a squashed data set that has the same structure as the original data with the addition of a weight for each pseudo data point that reflects the distribution of the original data into the initial groups. Any ML or SM method that accepts weights can be used to analyze the weighted pseudo data. By construction the resulting analyses will mimic the corresponding analyses on the original data set. Squashing should appeal to many of the sub-disciplines of