Lazy preservation: reconstructing websites by crawling the crawlers

Lazy preservation: reconstructing websites by crawling the crawlers
复制标题

懒惰保存:通过爬虫爬取重建网站

DOI:
10.1145/1183550.1183564
复制
发表时间:
2006
期刊:
影响因子:
3.7
通讯作者:
Michael L. Nelson
Michael L. Nelson
中科院分区:
医学3区
文献类型:
--
作者:
F. McCown;Joan A. Smith;Michael L. Nelson

文献摘要

被引文献

相似文献

通常直到网站或其网站管理员发生灾难性事件后才考虑网站备份。我们引入“惰性保存”——由于网络爬虫和缓存正常运行而执行的数字保存。惰性保存特别适合第三方;例如,一位老师重建了以前课程中使用的丢失的网站。我们通过使用网络存储库爬虫 Warrick 重建 24 个不同大小和组成的网站来评估惰性保存的有效性。由于任一存储库的完整性水平各不相同,我们的重建样本来自四个不同的网络存储库:Google (44%)、MSN (30%)、Internet Archive (19%) 和 Yahoo (7%)。我们还测量了发现和缓存 Web 资源所需的时间(10-103 天)以及删除后它们在缓存中保留的时间(7-61 天)。
Backup of websites is often not considered until after a catastrophic event has occurred to either the website or its webmaster. We introduce "lazy preservation" -- digital preservation performed as a result of the normal operation of web crawlers and caches. Lazy preservation is especially suitable for third parties; for example, a teacher reconstructing a missing website used in previous classes. We evaluate the effectiveness of lazy preservation by reconstructing 24 websites of varying sizes and composition using Warrick, a web-repository crawler. Because of varying levels of completeness in any one repository, our reconstructions sampled from four different web repositories: Google (44%), MSN (30%), Internet Archive (19%) and Yahoo (7%). We also measured the time required for web resources to be discovered and cached (10-103 days) as well as how long they remained in cache after deletion (7-61 days).