Discovery of Frequent Tree Structured Patterns in Semistructured Web Documents
Discovery of Frequent Tree Structured Patterns in Semistructured Web Documents
复制标题
DOI:
10.1007/3-540-45357-1_8
复制
发表时间:
2001-04
期刊:
影响因子:
--
通讯作者:
T. Miyahara;Takayoshi Shoudai;Tomoyuki Uchida;Kenichi Takahashi;H. Ueda
中科院分区:
文献类型:
--
作者:
T. Miyahara;Takayoshi Shoudai;Tomoyuki Uchida;Kenichi Takahashi;H. Ueda
Many documents such as Web documents or XML files have no rigid structure. Such semistructured documents have been rapidly increasing. We propose a new method for discovering frequent tree structured patterns in semistructured Web documents. We consider the data mining problem of finding all maximally frequent tag tree patterns in semistructured data such as Web documents. A tag tree pattern is an edge labeled tree which has hyperedges as variables. An edge label is a tag or a keyword in Web documents, and a variable can be substituted by any tree. So a tag tree pattern is suited for representing tree structured patterns in semistructured Web documents. We present an algorithm for finding all maximally frequent tag tree patterns. Also we report some experimental results on XML documents by using our algorithm.