NSDF-Catalog: Lightweight Indexing Service for Democratizing Data Delivery

NSDF-Catalog: Lightweight Indexing Service for Democratizing Data Delivery
复制标题

DOI:
10.1109/ucc56403.2022.00011
复制
发表时间:
2022-12
期刊:
2022 IEEE/ACM 15th International Conference on Utility and Cloud Computing (UCC)
影响因子:
--
通讯作者:
Jakob Luettgau;Christine R. Kirkpatrick;G. Scorzelli;Valerio Pascucci;G. Tarcea
Jakob Luettgau;Christine R. Kirkpatrick;G. Scorzelli;Valerio Pascucci;G. Tarcea
中科院分区:
其他
文献类型:
--
作者:
Jakob Luettgau;Christine R. Kirkpatrick;G. Scorzelli;Valerio Pascucci;G. Tarcea

文献摘要

被引文献

相似文献

在整个领域中生成了大量的科学数据。由于信息的大量信息,数据可发现性是一个挑战,特别是对于尚未生成数据或来自其他领域的科学家而言。作为NSF资助的国家科学数据结构(NSDF)倡议的一部分,我们开发了一个测试台,以证明可以克服数据可发现性的这些界限。为了支持这项工作,我们确定了跨科学领域索引大量科学数据的需求。我们提出了NSDF-catalog,这是一种用最小的元数据进行轻巧的索引服务,可以补充现有的域特异性和丰富的 - 米达塔(Metadata Col)。 NSDF-catalog旨在促进灵活的微服务中的多个相关目标,以:(i)在NSDF Federation内的原始存储库中协调数据运动和复制; (ii)建立现有科学数据的清单,以告知下一代网络基础设施的设计; (iii)提供了一套用于发现跨学科研究数据集的工具。我们的服务在文件或对象级别上以细粒度为科学数据,以告知数据分发策略并从消费者的角度改善用户的体验,目的是允许端到端的数据流优化。
Across domains massive amounts of scientific data are generated. Because of the large volume of information, data discoverability is a challenge, especially for scientists who have not generated the data or are from other domains. As part of the NSF-funded National Science Data Fabric (NSDF) initiative, we developed a testbed to demonstrate that these boundaries to data discoverability can be overcome. In support of this effort, we identify the need for indexing large-amounts of scientific data across scientific domains. We propose NSDF-Catalog, a lightweight indexing service with minimal metadata that complements existing domain-specific and rich-metadata col-lections. NSDF-Catalog is designed to facilitate multiple related objectives within a flexible microservice to: (i) coordinate data movements and replication of data from origin repositories within the NSDF federation; (ii) build an inventory of existing scientific data to inform the design of next-generation cyberinfrastructure; and (iii) provide a suite of tools for discovery of datasets for cross-disciplinary research. Our service indexes scientific data at a fine-granularity at the file or object level to inform data distribution strategies and to improve the experience for users from the consumer perspective, with the goal of allowing end-to-end dataflow optimizations.