Web page classification without the web page
Web page classification without the web page
复制标题
没有网页的网页分类
DOI:
10.1145/1013367.1013426
复制
发表时间:
2004
期刊:
影响因子:
--
通讯作者:
Min
中科院分区:
文献类型:
--
作者:
Min
Uniform resource locators (URLs), which mark the address of a resource on the World Wide Web, are often human-readable and can hint at the category of the resource. This paper explores the use of URLs for webpage categorization via a two-phase pipeline of word segmentation/expansion and classification. We quantify its performance against document-based methods, which require the retrieval of the source document.