Web unit mining: finding and classifying subgraphs of web pages

Web unit mining: finding and classifying subgraphs of web pages
复制标题

Web单元挖掘:查找网页的子图并进行分类

DOI:
10.1145/956863.956885
复制
发表时间:
2003
期刊:
The Library Quarterly
影响因子:
--
通讯作者:
Ee
Ee
中科院分区:
--
文献类型:
--
作者:
Aixin Sun;Ee

文献摘要

被引文献

相似文献

在网页分类中,大多数研究人员假设要分类的对象是来自一个或多个网站的各个网页。在实践中,该假设过于严格,因为网页本身可能并不总是对应于给予分类任务的某些语义概念(或类别)的概念实例。在本文中,我们希望放宽这一假设,并允许概念实例由网页的子图或一组网页来表示。我们确定了删除假设后需要解决的几个新问题,并制定了网络单元挖掘问题。我们还提出了一种迭代网络单元挖掘(iWUM)方法,该方法首先使用有关网站结构的一些知识来查找网页的子图。从这些网络子图中,网络单元被构建并以迭代方式分类为语义概念(或类别)。我们使用 WebKB 数据集进行的实验表明,iWUM 提高了整体分类性能,并且在网站的结构化部分上效果很好。
In web classification, most researchers assume that the objects to classify are individual web pages from one or more web sites. In practice, the assumption is too restrictive since a web page itself may not always correspond to a concept instance of some semantic concept (or category) given to the classification task. In this paper, we want to relax this assumption and allow a concept instance to be represented by a subgraph of web pages or a set of web pages. We identify several new issues to be addressed when the assumption is removed, and formulate the web unit mining problem. We also propose an iterative web unit mining (iWUM) method that first finds subgraphs of web pages using some knowledge about web site structure. From these web subgraphs, web units are constructed and classified into semantic concepts (or categories) in an iterative manner. Our experiments using the WebKB dataset showed that iWUM improves the overall classification performance and works very well on the more structured parts of a web site.