Focused crawling: a new approach to topic-specific Web resource discovery

Focused crawling: a new approach to topic-specific Web resource discovery
复制标题

DOI:
10.1016/s1389-1286(99)00052-3
复制
发表时间:
1999-05-17
期刊:
COMPUTER NETWORKS-THE INTERNATIONAL JOURNAL OF COMPUTER AND TELECOMMUNICATIONS NETWORKING
影响因子:
--
通讯作者:
Dom, B
Dom, B
中科院分区:
其他
文献类型:
--
作者:
Chakrabarti, S;van den Berg, M;Dom, B

文献摘要

被引文献

相似文献

万维网的快速增长给通用爬虫和搜索引擎带来了前所未有的扩展挑战。在本文中,我们描述了一种新的超文本资源发现系统,称为聚焦爬虫。专注的爬虫的目标是有选择地寻找与预定义主题集相关的页面。主题不是使用关键字指定的,而是使用示例文档指定的。专注的爬行程序不是收集所有可访问的 Web 文档并对其建立索引以便能够回答所有可能的即席查询,而是分析其爬行边界以查找可能与爬行最相关的链接,并避开 Web 中不相关的区域。这可以显着节省硬件和网络资源,并有助于使爬行保持最新状态。为了实现这种目标导向的爬行,我们设计了两个超文本挖掘程序来指导我们的爬虫:一个分类器,用于评估超文本文档与焦点主题的相关性;一个提取器,用于识别超文本节点,这些节点是通过几个链接访问许多相关页面的绝佳访问点。我们报告了使用不同特异性级别的多个主题进行的广泛的集中爬行实验。集中爬行可以稳定地获取相关页面,而标准爬行很快就会迷失方向,即使它们是从相同的根集开始的。集中爬行对于起始 URL 集中的大扰动具有鲁棒性。尽管存在这些扰动,它仍然发现了大部分重叠的资源集。它还能够探索和发现距起始集数十个链接的有价值的资源,同时仔细修剪可能位于同一半径内的数百万个页面。我们的轶事表明,使用适度的桌面硬件,集中爬行对于构建特定主题的高质量 Web 文档集合非常有效。 (C) 1999 年由 Elsevier Science B.V. AU 出版,保留权利。
The rapid growth of the World-Wide Web poses unprecedented scaling challenges for general-purpose crawlers and search engines. In this paper we describe a new hypertext resource discovery system called a Focused Crawler. The goal of a focused crawler is to selectively seek out pages that are relevant to a pre-defined set of topics. The topics are specified not using keywords, but using exemplary documents. Rather than collecting and indexing all accessible Web documents to be able to answer all possible ad-hoc queries, a focused crawler analyzes its crawl boundary to find the links that are likely to be most relevant for the crawl, and avoids irrelevant regions of the Web. This leads to significant savings in hardware and network resources, and helps keep the crawl more up-to-date.To achieve such goal-directed crawling, we designed two hypertext mining programs that guide our crawler: a classifier that evaluates the relevance of a hypertext document with respect to the focus topics, and a distiller that identifies hypertext nodes that are great access points to many relevant pages within a few links. We report on extensive focused-crawling experiments using several topics at different levels of specificity. Focused crawling acquires relevant pages steadily while standard crawling quickly loses its way, even though they are started from the same root set. Focused crawling is robust against large perturbations in the starting set of URLs. It discovers largely overlapping sets of resources in spite of these perturbations. It is also capable of exploring out and discovering valuable resources that are dozens of links away from the start set, while carefully pruning the millions of pages that may lie within this same radius. Our anecdotes suggest that focused crawling is very effective for building high-quality collections of Web documents on specific topics, using modest desktop hardware. (C) 1999 Published by Elsevier Science B.V. AU rights reserved.