Domain-specific Web site identification: the CROSSMARC focused Web crawler

Domain-specific Web site identification: the CROSSMARC focused Web crawler
复制标题

特定领域的网站识别:专注于 CROSSMARC 的网络爬虫

DOI:
--
复制
发表时间:
2003
期刊:
--
影响因子:
--
通讯作者:
Shipra Dingare
Shipra Dingare
中科院分区:
--
文献类型:
--
作者:
K. Stamatakis;V. Karkaletsis;G. Paliouras;James Horlock;Claire Grover;J. Curran;Shipra Dingare

文献摘要

被引文献

相似文献

本文介绍了用于识别特定领域的网站,已实施的EC资助的R&D项目,CROSSMARC的一部分,技术。该项目旨在开发从特定领域网页中提取有趣信息的技术。因此,重要的是CROSSMARC识别网站,其中感兴趣的领域特定的网页驻留(重点网络爬行)。这就是CROSSMARC网络爬虫的作用。
This paper presents techniques for identifying domain specific web sites that have been implemented as part of the EC-funded R&D project, CROSSMARC. The project aims to develop technology for extracting interesting information from domain-specific web pages. It is therefore important for CROSSMARC to identify web sites in which interesting domain specific pages reside (focused web crawling). This is the role of the CROSSMARC web crawler.