Efficient Substructure Discovery from Large Semi-Structured Data

Efficient Substructure Discovery from Large Semi-Structured Data
复制标题

DOI:
10.1137/1.9781611972726.10
复制
发表时间:
2001-10
期刊:
IEICE Trans. Inf. Syst.
影响因子:
--
通讯作者:
Tatsuya Asai;K. Abe;Shinji Kawasoe;H. Sakamoto;Hiroki Arimura;S. Arikawa
Tatsuya Asai;K. Abe;Shinji Kawasoe;H. Sakamoto;Hiroki Arimura;S. Arikawa
中科院分区:
其他
文献类型:
--
作者:
Tatsuya Asai;K. Abe;Shinji Kawasoe;H. Sakamoto;Hiroki Arimura;S. Arikawa

文献摘要

被引文献

相似文献

本文研究了半结构化数据的数据挖掘问题。将半结构化数据建模为带标签的有序树,提出了一种从大量半结构化数据中发现频繁子结构的高效算法。通过扩展Bayardo(SIGMOD‘98)开发的发现长项集的枚举技术,我们的算法几乎线性地扩展了输入集合中包含的最大树模式的总大小,轻微地依赖于最长模式的大小。我们还开发了几种修剪技术,大大加快了搜索速度。在Web数据上的实验表明,该算法结合所提出的剪枝技术,能够在较宽的参数范围内有效地运行在真实数据集上。
In this paper, we consider a data mining problem for semi-structured data. Modeling semi-structured data as labeled ordered trees, we present an efficient algorithm for discovering frequent substructures from a large collection of semi-structured data. By extending the enumeration technique developed by Bayardo (SIGMOD'98) for discovering long itemsets, our algorithm scales almost linearly in the total size of maximal tree patterns contained in an input collection depending mildly on the size of the longest pattern. We also developed several pruning techniques that significantly speed-up the search. Experiments on Web data show that our algorithm runs efficiently on real-life datasets combined with proposed pruning techniques in the wide range of parameters.