Incremental Crawling with Heritrix

Incremental Crawling with Heritrix
复制标题

使用 Heritrix 进行增量爬行

DOI:
--
复制
发表时间:
2010
期刊:
--
影响因子:
--
通讯作者:
Kristinn Sigurðsson
Kristinn Sigurðsson
中科院分区:
--
文献类型:
--
作者:
Kristinn Sigurðsson

文献摘要

被引文献

相似文献

Heritrix网络爬虫的目标是成为世界上第一个开源的、可扩展的、网络规模的、具有档案质量的网络爬虫。然而,它的爬行策略仅限于快照爬行。本文报告了在其功能的基础上添加增量爬行功能的工作。我们首先讨论与快照爬行相对的增量爬行的概念,然后讨论设计有效的增量策略的可能方法。对我们所做的实现进行了概述,并讨论了其局限性和优点。然后,我们报告了新软件的初步试验结果,这些结果运行良好。最后,我们讨论了仍未解决的问题和未来可能的改进。
The Heritrix web crawler aims to be the world's first open source, extensible, web-scale, archival-quality web crawler. It has however been limited in its crawling strategies to snapshot crawling. This paper reports on work to add the ability to conduct incremental crawls to its capabilities. We first discuss the concept of incremental crawling as opposed to snapshot crawling and then the possible ways to design an effective incremental strategy. An overview is given of the implementation that we did, its limits and strengths are discussed. We then report on the results of initial experimentation with the new software which have gone well. Finally, we discuss issues that remain unresolved and possible future improvements.