Breadth-first crawling yields high-quality pages

Breadth-first crawling yields high-quality pages
复制标题

DOI:
10.1145/371920.371965
复制
发表时间:
2001-05
期刊:
影响因子:
3.9
通讯作者:
Marc Najork;J. Wiener
Marc Najork;J. Wiener
中科院分区:
计算机科学3区
文献类型:
--
作者:
Marc Najork;J. Wiener

文献摘要

被引文献

相似文献

本文研究了在对3.28亿个唯一页面进行网络抓取时,随着时间的推移下载的页面的平均页面质量。我们使用基于连接性的PageRank度量来衡量页面的质量。我们表明,在广度优先搜索顺序遍历网络图是一个很好的爬行策略,因为它往往会发现高质量的网页在抓取早期。
This paper examines the average page quality over time of pages downloaded during a web crawl of 328 million unique pages. We use the connectivity-based metric PageRank to measure the quality of a page. We show that traversing the web graph in breadth-first search order is a good crawling strategy, as it tends to discover high-quality pages early on in the crawl.