Comparing sets of patterns with the Jaccard index

Comparing sets of patterns with the Jaccard index
复制标题

将模式集与 Jaccard 指数进行比较

DOI:
--
复制
发表时间:
2018
影响因子:
2.1
通讯作者:
M. Islam
M. Islam
中科院分区:
--
文献类型:
--
作者:
Sam Fletcher;M. Islam

文献摘要

被引文献

相似文献

从数据中提取知识的能力从一开始就是数据挖掘的驱动力,甚至在此之前很久就是统计建模的驱动力。可操作的知识通常采用模式的形式,其中可以使用一组先行词来推断结果。在本文中,我们提供了一个比较不同模式集问题的解决方案。我们的解决方案允许比较来自不同技术(例如不同分类算法)的模式集,或者来自不同数据样本(例如时间数据或由于隐私原因而受到干扰的数据)的模式集。我们建议使用Jaccard索引通过将每个模式转换为集合中的单个元素来度量模式集之间的相似性。我们的测量侧重于提供概念简单性、计算简单性、可解释性和广泛的适用性。将此度量的结果与实际数据挖掘场景中的预测精度进行比较。
The ability to extract knowledge from data has been the driving force of Data Mining since its inception, and of statistical modeling long before even that. Actionable knowledge often takes the form of patterns, where a set of antecedents can be used to infer a consequent. In this paper we offer a solution to the problem of comparing different sets of patterns. Our solution allows comparisons between sets of patterns that were derived from different techniques (such as different classification algorithms), or made from different samples of data (such as temporal data or data perturbed for privacy reasons). We propose using the Jaccard index to measure the similarity between sets of patterns by converting each pattern into a single element within the set. Our measure focuses on providing conceptual simplicity, computational simplicity, interpretability, and wide applicability. The results of this measure are compared to prediction accuracy in the context of a real-world data mining scenario.