Efficient Distribution Mining and Classification

Efficient Distribution Mining and Classification
复制标题

高效分布挖掘与分类

DOI:
10.1137/1.9781611972788.58
复制
发表时间:
2008
期刊:
Dermatologic surgery : official publication for American Society for Dermatologic Surgery [et al.]
影响因子:
--
通讯作者:
C. Faloutsos
C. Faloutsos
中科院分区:
--
文献类型:
--
作者:
Yasushi Sakurai;Rosalynn Chong;Lei Li;C. Faloutsos

文献摘要

被引文献

相似文献

我们定义并解决了“分布分类”问题,以及一般的“分布挖掘”问题。给定n个分布(即,云),我们希望将它们分为k类,以找到模式,规则和离群点云。例如,考虑物品销售的二维情况,其中,对于每个售出的物品,我们记录单价和数量;然后,每个客户被表示为二维点的分布/云(他购买的每个物品一个)。我们希望将相似的用户分组在一起,例如,用于市场细分、异常/欺诈检测。我们建议D-Mine来实现这一目标。我们的主要贡献是定理3.1,它展示了如何使用小波来加速云相似性计算。在合成和真实的多维数据集上进行的大量实验表明,我们的方法比简单的实现方法快了400倍,并且具有相当的(偶尔更好的)分类质量。
We define and solve the problem of “distribution classifi-cation”, and, in general, “distribution mining”. Given n distributions (i.e., clouds) of multi-dimensional points, we want to classify them into k classes, to find patterns, rules and out-lier clouds. For example, consider the 2-d case of sales of items, where, for each item sold, we record the unit price and quantity; then, each customer is represented as a distribution/cloud of 2-d points (one for each item he bought). We want to group similar users together, e.g., for market segmentation, anomaly/fraud detection. We propose D-Mine to achieve this goal. Our main contribution is Theorem 3.1, which shows how to use wavelets to speed up the cloud-similarity computations. Extensive experiments on both synthetic and real multi-dimensional data sets show that our method achieves up to 400 faster wall-clock time over the naive implementation, with comparable (and occasionally better) classification quality.