Lazy preservation: reconstructing websites by crawling the crawlers
Lazy preservation: reconstructing websites by crawling the crawlers
复制标题
懒惰保存:通过爬虫爬取重建网站
DOI:
10.1145/1183550.1183564
复制
发表时间:
2006
期刊:
影响因子:
3.7
通讯作者:
Michael L. Nelson
中科院分区:
文献类型:
--
作者:
F. McCown;Joan A. Smith;Michael L. Nelson
Backup of websites is often not considered until after a catastrophic event has occurred to either the website or its webmaster. We introduce "lazy preservation" -- digital preservation performed as a result of the normal operation of web crawlers and caches. Lazy preservation is especially suitable for third parties; for example, a teacher reconstructing a missing website used in previous classes. We evaluate the effectiveness of lazy preservation by reconstructing 24 websites of varying sizes and composition using Warrick, a web-repository crawler. Because of varying levels of completeness in any one repository, our reconstructions sampled from four different web repositories: Google (44%), MSN (30%), Internet Archive (19%) and Yahoo (7%). We also measured the time required for web resources to be discovered and cached (10-103 days) as well as how long they remained in cache after deletion (7-61 days).