A Serverless Framework for Distributed Bulk Metadata Extraction

A Serverless Framework for Distributed Bulk Metadata Extraction
复制标题

用于分布式批量元数据提取的无服务器框架

DOI:
10.1145/3431379.3460636
复制
发表时间:
2021
期刊:
Proceedings of the 30th International Symposium on High-Performance Parallel and Distributed Computing (HPDC
影响因子:
--
通讯作者:
Foster, Ian
Foster, Ian
中科院分区:
--
文献类型:
--
作者:
Skluzacek, Tyler J.;Wong, Ryan;Li, Zhuozhao;Chard, Ryan;Chard, Kyle;Foster, Ian

文献摘要

参考文献

被引文献

相似文献

我们介绍Xtract,一个自动化和可扩展的系统,用于从大型分布式研究数据存储库中提取批量元数据。Xtract将元数据提取器的应用程序编排到文件组,确定将哪些提取器应用于每个文件,以及对于每个提取器和文件,在何处执行。基于funcX联邦FaaS平台构建的混合计算模型使Xtract能够通过将每个提取任务分配到最合适的位置来平衡提取时间和数据传输成本之间的权衡。在一系列云和超级计算机上的实验表明,Xtract可以通过在数千个节点上协调基于容器的提取器的并发执行来有效地处理数百万个文件存储库。我们通过将Xtract应用于大型半策划的科学数据存储库和未经策划的科学Google Drive存储库来突出Xtract的灵活性。我们表明,通过在分散的存储和计算节点上远程编排元数据提取,Xtract可以在50%的时间内处理大型存储库,只需将相同的数据传输到同一计算设施内的机器。我们还表明,当传输数据是必要的(例如,没有本地计算可用),Xtract可以扩展到处理文件的速度与接收文件的速度一样快,即使是在多GB/s的网络上。
We introduce Xtract, an automated and scalable system for bulk metadata extraction from large, distributed research data repositories. Xtract orchestrates the application of metadata extractors to groups of files, determining which extractors to apply to each file and, for each extractor and file, where to execute. A hybrid computing model, built on the funcX federated FaaS platform, enables Xtract to balance tradeoffs between extraction time and data transfer costs by dispatching each extraction task to the most appropriate location. Experiments on a range of clouds and supercomputers show that Xtract can efficiently process multi-million-file repositories by orchestrating the concurrent execution of container-based extractors on thousands of nodes. We highlight the flexibility of Xtract by applying it to a large, semi-curated scientific data repository and to an uncurated scientific Google Drive repository. We show that by remotely orchestrating metadata extraction across decentralized storage and compute nodes, Xtract can process large repositories in 50% of the time it takes just to transfer the same data to a machine within the same computing facility. We also show that when transferring data is necessary (e.g., no local compute is available), Xtract can scale to process files as fast as they are received, even over a multi-GB/s network.
DOI: 10.1145/3219104.3219159
发表时间: 2018-07
期刊: Proceedings of the Practice and Experience on Advanced Research Computing
影响因子: --
作者:
Luigi Marini;I. Gutierrez-Polo;R. Kooper;Sandeep Puthanveetil Satheesan;M. Burnette;J. Lee;Todd Nicholson;Yan Zhao;Kenton McHenry
通讯作者: Luigi Marini;I. Gutierrez-Polo;R. Kooper;Sandeep Puthanveetil Satheesan;M. Burnette;J. Lee;Todd Nicholson;Yan Zhao;Kenton McHenry
衡量元数据质量
DOI: --
发表时间: 2019
期刊:
影响因子: --
作者:
P. Király
通讯作者: P. Király
光源科学大数据远程访问接口
DOI: 10.1109/bdc.2015.37
发表时间: 2015
期刊: 2015 IEEE/ACM 2nd International Symposium on Big Data Computing (BDC)
影响因子: --
作者:
J. Wozniak;K. Chard;B. Blaiszik;Ray Osborn;M. Wilde;Ian T Foster
通讯作者: Ian T Foster