Managing duplicates across sequential crawls

Managing duplicates across sequential crawls
复制标题

跨顺序爬网管理重复项

DOI:
--
复制
发表时间:
2010
期刊:
--
影响因子:
--
通讯作者:
Kristinn Sigurðsson
Kristinn Sigurðsson
中科院分区:
--
文献类型:
--
作者:
Kristinn Sigurðsson

文献摘要

被引文献

相似文献

对于许多对万维网进行归档的组织来说,处理在收获轮次之间保持不变的文档是一个重要的主题。本文讨论了处理此问题的一些关键问题,然后概述了一种简单而有效的方法来管理至少一部分问题。这是通过流行的网络爬虫 Heritrix 的附加模块的形式完成的。该论文包含使用这个新软件的爬行结果。最后,讨论了局限性以及改进重复处理所需的一些未来工作。
Dealing with documents that remain unchanged between harvesting rounds is an important topic for many organizations archiving the World Wide Web. This paper discusses some of the key problems in dealing with this and then outlines a simple, yet effective way of managing at least a part of it. This is done in form of an add-on module for the popular web crawler Heritrix. The paper contains the results of crawls using this new software. Finally, there is a discussion on the limitations and some of the future work needed to improve duplicate handling.