Discovery of Frequent Tag Tree Patterns in Semistructured Web Documents

Discovery of Frequent Tag Tree Patterns in Semistructured Web Documents
复制标题

DOI:
10.1007/3-540-47887-6_35
复制
发表时间:
2002-05
期刊:
--
影响因子:
--
通讯作者:
T. Miyahara;Yusuke Suzuki;Takayoshi Shoudai;Tomoyuki Uchida;Kenichi Takahashi;H. Ueda
T. Miyahara;Yusuke Suzuki;Takayoshi Shoudai;Tomoyuki Uchida;Kenichi Takahashi;H. Ueda
中科院分区:
其他
文献类型:
--
作者:
T. Miyahara;Yusuke Suzuki;Takayoshi Shoudai;Tomoyuki Uchida;Kenichi Takahashi;H. Ueda

文献摘要

被引文献

相似文献

许多Web文档(例如HTML文件和XML文件)没有严格的结构,被称为半结构化数据。一般来说,这种半结构化 Web 文档由具有有序子项的有根树表示。我们提出了一种新方法,通过使用标签树模式作为假设来发现半结构化 Web 文档中的频繁树结构模式。标签树模式是具有结构化变量的有序子节点的边缘标记树。边缘标签是此类Web文档中的标签或关键字,并且变量可以被任意树替代。因此,标签树模式适合表示此类 Web 文档中的树结构模式。首先,我们表明很难计算最佳频繁标签树模式。因此我们提出了一种生成所有最大频繁标签树模式的算法并给出了它的正确性。最后,我们报告了我们算法的一些实验结果。尽管该算法效率不高,但实验表明我们可以提取这些数据中的特征树结构模式。
Many Web documents such as HTML files and XML files have no rigid structure and are called semistructured data. In general, such semistructured Web documents are represented by rooted trees with ordered children. We propose a new method for discovering frequent tree structured patterns in semistructured Web documents by using a tag tree pattern as a hypothesis. A tag tree pattern is an edge labeled tree with ordered children which has structured variables. An edge label is a tag or a keyword in such Web documents, and a variable can be substituted by an arbitrary tree. So a tag tree pattern is suited for representing tree structured patterns in such Web documents. First we show that it is hard to compute the optimum frequent tag tree pattern. So we present an algorithm for generating all maximally frequent tag tree patterns and give the correctness of it. Finally, we report some experimental results on our algorithm. Although this algorithm is not efficient, experiments show that we can extract characteristic tree structured patterns in those data.