Web page classification without the web page

Web page classification without the web page
复制标题

没有网页的网页分类

DOI:
10.1145/1013367.1013426
复制
发表时间:
2004
期刊:
WWW Alt. '04
影响因子:
--
通讯作者:
Min
Min
中科院分区:
--
文献类型:
--
作者:
Min

文献摘要

被引文献

相似文献

统一资源定位符(URL),它标记了万维网上的资源地址,通常是人类可读的,并且可以暗示资源的类别。本文探讨了使用URL的网页分类通过两个阶段的管道分词/扩展和分类。我们量化其性能对基于文档的方法,这需要检索的源文件。
Uniform resource locators (URLs), which mark the address of a resource on the World Wide Web, are often human-readable and can hint at the category of the resource. This paper explores the use of URLs for webpage categorization via a two-phase pipeline of word segmentation/expansion and classification. We quantify its performance against document-based methods, which require the retrieval of the source document.