Summarizing itemset patterns: a profile-based approach

Summarizing itemset patterns: a profile-based approach
复制标题

DOI:
10.1145/1081870.1081907
复制
发表时间:
2005-08
期刊:
--
影响因子:
--
通讯作者:
Xifeng Yan;Hong Cheng;Jiawei Han;Dong Xin
Xifeng Yan;Hong Cheng;Jiawei Han;Dong Xin
中科院分区:
其他
文献类型:
--
作者:
Xifeng Yan;Hong Cheng;Jiawei Han;Dong Xin

文献摘要

被引文献

相似文献

频繁模式挖掘已经被广泛地研究了可扩展的方法,用于挖掘各种类型的模式,包括项集,序列和图。然而,频繁模式挖掘的瓶颈不是效率,而是在可解释性,由于挖掘过程中产生的模式的巨大数量。在本文中,我们研究了如何总结一个项目集模式的集合使用只有K个代表,少量的模式,用户可以很容易地处理。K代表不仅应该覆盖大多数频繁模式,而且应该近似它们的支持。建立了一个生成模型来提取和描述这些代表,在这个模型下,模式的支持度可以很容易地恢复,而不需要咨询原始数据集。基于恢复误差,我们提出了一个质量度量函数来确定参数K的最佳值。多项式时间算法与几种优化算法一起开发,以提高效率。实验研究表明,我们可以获得紧凑的摘要在真实的数据集。
Frequent-pattern mining has been studied extensively on scalable methods for mining various kinds of patterns including itemsets, sequences, and graphs. However, the bottleneck of frequent-pattern mining is not at the efficiency but at the interpretability, due to the huge number of patterns generated by the mining process.In this paper, we examine how to summarize a collection of itemset patterns using only K representatives, a small number of patterns that a user can handle easily. The K representatives should not only cover most of the frequent patterns but also approximate their supports. A generative model is built to extract and profile these representatives, under which the supports of the patterns can be easily recovered without consulting the original dataset. Based on the restoration error, we propose a quality measure function to determine the optimal value of parameter K. Polynomial time algorithms are developed together with several optimization heuristics for efficiency improvement. Empirical studies indicate that we can obtain compact summarization in real datasets.