Managing duplicates across sequential crawls
Managing duplicates across sequential crawls
复制标题
跨顺序爬网管理重复项
DOI:
--
复制
发表时间:
2010
期刊:
影响因子:
--
通讯作者:
Kristinn Sigurðsson
中科院分区:
文献类型:
--
作者:
Kristinn Sigurðsson
Dealing with documents that remain unchanged between harvesting rounds is an important topic for many organizations archiving the World Wide Web. This paper discusses some of the key problems in dealing with this and then outlines a simple, yet effective way of managing at least a part of it. This is done in form of an add-on module for the popular web crawler Heritrix. The paper contains the results of crawls using this new software. Finally, there is a discussion on the limitations and some of the future work needed to improve duplicate handling.