Discovery of Frequent Tree Structured Patterns in Semistructured Web Documents

Discovery of Frequent Tree Structured Patterns in Semistructured Web Documents
复制标题

DOI:
10.1007/3-540-45357-1_8
复制
发表时间:
2001-04
期刊:
--
影响因子:
--
通讯作者:
T. Miyahara;Takayoshi Shoudai;Tomoyuki Uchida;Kenichi Takahashi;H. Ueda
T. Miyahara;Takayoshi Shoudai;Tomoyuki Uchida;Kenichi Takahashi;H. Ueda
中科院分区:
其他
文献类型:
--
作者:
T. Miyahara;Takayoshi Shoudai;Tomoyuki Uchida;Kenichi Takahashi;H. Ueda

文献摘要

被引文献

相似文献

许多文档,如Web文档或XML文件,都没有严格的结构。这样的半结构化文档一直在迅速增加。提出了一种在半结构化Web文档中发现频繁树结构模式的新方法。我们考虑了在Web文档等半结构化数据中发现所有最大频繁标签树模式的数据挖掘问题。标签树模式是以超边为变量的边标记树。边缘标签是Web文档中的标签或关键字,变量可以由任何树替换。因此,标签树模式适合于表示半结构化Web文档中的树形结构模式。提出了一种寻找所有最大频繁标签树模式的算法。文中还给出了我们的算法在XML文档上的实验结果。
Many documents such as Web documents or XML files have no rigid structure. Such semistructured documents have been rapidly increasing. We propose a new method for discovering frequent tree structured patterns in semistructured Web documents. We consider the data mining problem of finding all maximally frequent tag tree patterns in semistructured data such as Web documents. A tag tree pattern is an edge labeled tree which has hyperedges as variables. An edge label is a tag or a keyword in Web documents, and a variable can be substituted by any tree. So a tag tree pattern is suited for representing tree structured patterns in semistructured Web documents. We present an algorithm for finding all maximally frequent tag tree patterns. Also we report some experimental results on XML documents by using our algorithm.