Cloud-Native Repositories for Big Scientific Data

Cloud-Native Repositories for Big Scientific Data
复制标题

DOI:
10.1109/mcse.2021.3059437
复制
发表时间:
2021-03-01
影响因子:
2.1
通讯作者:
Signell, Richard P.
Signell, Richard P.
中科院分区:
计算机科学4区
文献类型:
--
作者:
Abernathey, Ryan P.;Blackmon-Luca, Charles C.;Signell, Richard P.

文献摘要

被引文献

相似文献

科学数据传统上是通过从数据服务器下载到本地计算机来分发的。随着科学数据集向PB级增长,这种工作方式受到限制。本文中定义的“云原生数据存储库”提供了几个优于传统数据存储库的优势性能、可靠性、成本效益、协作、可再现性、创造性、下游影响以及访问和包含性。这些目标激发了一组云原生数据存储库的最佳实践:分析就绪数据、云优化(阿科)格式以及与数据近似计算的松散耦合。Pangeo项目通过使用开源科学Python工具开发了这些原则的原型实现。通过提供阿科数据目录以及按需、可扩展的分布式计算,Pangeo使用户能够以超过10 GB/s的速度处理大数据。为了实现云计算在科学研究方面的全部潜力,必须解决几个挑战,例如组织资金、培训用户和执行数据隐私要求。
Scientific data have traditionally been distributed via downloads from data server to local computer. This way of working suffers from limitations as scientific datasets grow toward the petabyte scale. A "cloud-native data repository," as defined in this article, offers several advantages over traditional data repositories-performance, reliability, cost-effectiveness, collaboration, reproducibility, creativity, downstream impacts, and access and inclusion. These objectives motivate a set of best practices for cloud-native data repositories: analysis-ready data, cloud-optimized (ARCO) formats, and loose coupling with data-proximate computing. The Pangeo Project has developed a prototype implementation of these principles by using open-source scientific Python tools. By providing an ARCO data catalog together with on-demand, scalable distributed computing, Pangeo enables users to process big data at rates exceeding 10 GB/s. Several challenges must be resolved in order to realize cloud computing's full potential for scientific research, such as organizing funding, training users, and enforcing data privacy requirements.