ArchiveSpark: Efficient Web archive access, extraction and derivation

ArchiveSpark: Efficient Web archive access, extraction and derivation
复制标题

DOI:
10.1145/2910896.2910902
复制
发表时间:
2016-06
期刊:
2016 IEEE/ACM Joint Conference on Digital Libraries (JCDL)
影响因子:
--
通讯作者:
Helge Holzmann;V. Goel;Avishek Anand
Helge Holzmann;V. Goel;Avishek Anand
中科院分区:
其他
文献类型:
--
作者:
Helge Holzmann;V. Goel;Avishek Anand

文献摘要

被引文献

相似文献

网络档案是各学科研究人员的宝贵资源。然而,使用它们作为学术来源,研究人员需要一个工具,提供有效的访问Web存档数据的提取和派生的较小的数据集。除了有效的访问,我们确定了其他五个目标的基础上,实际研究人员的需求,如易用性,可扩展性和可重用性。为了实现这些目标,我们提出了ArchiveSpark,一个高效的,分布式的Web档案处理框架,建立了一个研究语料库的工作,现有的和标准化的数据格式通常由Web存档机构。通过使用广泛可用的元数据索引,ArchiveSpark中的性能优化可以显著提高数据处理的速度。我们的基准测试表明,ArchiveSpark比其他方法更快,无需依赖任何额外的数据存储,同时通过将查询和派生与外部工具无缝集成来提高可用性。
Web archives are a valuable resource for researchers of various disciplines. However, to use them as a scholarly source, researchers require a tool that provides efficient access to Web archive data for extraction and derivation of smaller datasets. Besides efficient access we identify five other objectives based on practical researcher needs such as ease of use, extensibility and reusability. Towards these objectives we propose ArchiveSpark, a framework for efficient, distributed Web archive processing that builds a research corpus by working on existing and standardized data formats commonly held by Web archiving institutions. Performance optimizations in ArchiveSpark, facilitated by the use of a widely available metadata index, result in significant speed-ups of data processing. Our benchmarks show that ArchiveSpark is faster than alternative approaches without depending on any additional data stores while improving usability by seamlessly integrating queries and derivations with external tools.