Association-based similarity testing and its applications

Association-based similarity testing and its applications
复制标题

DOI:
10.3233/ida-2003-7304
复制
发表时间:
2003-08
期刊:
Intell. Data Anal.
影响因子:
--
通讯作者:
Tao Li;Mitsunori Ogihara;Shenghuo Zhu
Tao Li;Mitsunori Ogihara;Shenghuo Zhu
中科院分区:
其他
文献类型:
--
作者:
Tao Li;Mitsunori Ogihara;Shenghuo Zhu

文献摘要

被引文献

相似文献

本文提出了一种新的基于关联的篮子数据集之间的相似性度量。新的度量是使用受信息熵启发的公式从支持计数计算的。在真实的数据集和人工合成数据集上的实验表明了该方法的有效性。本文接着研究了相似性度量的应用。它首先研究了使用相似性度量在分类数据库属性集之间找到映射的问题。提出了一种通用的方法来确定这样的映射。该方法的实现基于本文提出的相似性度量,其性能进行了评估和验证。此外,本文还探讨了相似性度量在分布式数据挖掘中的应用。
This paper proposes a new similarity measure between basket datasets based on associations. The new measure is calculated from support counts using a formula inspired by information entropy. Experiments on both real and synthetic datasets show the effectiveness of the measure. This paper then investigates the applications of the similarity measure. It first studies the problem of finding a mapping between categorical database attribute sets using similarity measures. A generic approach for identifying such a mapping is proposed. The approach is implemented based on the similarity measure proposed in the paper and its performance has been evaluated and validated. Moreover, this paper also explores the applications of using the similarity measure to mine distributed datasets.