Tapis: An API Platform for Reproducible, Distributed Computational Research

Tapis: An API Platform for Reproducible, Distributed Computational Research
复制标题

Tapis:用于可重复的分布式计算研究的 API 平台

DOI:
10.1007/978-3-030-73100-7_61
复制
发表时间:
2021
影响因子:
2.7
通讯作者:
G. Jacobs
G. Jacobs
中科院分区:
工程技术3区
文献类型:
--
作者:
Joe Stubbs;Richard Cardone;Mike Packard;Anagha Jamthe;Smruti Padhy;Steve Terry;Julia Looney;Joseph Meiring;S. Black;M. Dahan;S. Cleveland;G. Jacobs

文献摘要

被引文献

相似文献

现代计算研究越来越多地跨越多个地理分布的数据中心,并利用仪器,实验设施和国家和区域网络基础设施(CI)网络。Tapis是一个开源的API平台,由德克萨斯大学奥斯汀分校的德克萨斯高级计算中心开发,用于提高分布式计算实验的可重复性并最大限度地缩短求解时间。Tapis的核心功能包括数据管理和代码执行,一个细粒度的权限系统,使对象能够被私人保存,与个人共享或“发布”到社区,以及出处端点,暴露Tapis在分析时收集的详细历史,使工作流程能够重复,结果能够重现。在本文中,我们描述了Tapis平台从2008年开始的演变,并讨论了该项目的发展和成功,以及导致2019年9月由美国国家科学基金会资助的新设计工作的挑战和限制。我们详细介绍了新系统,包括参考架构和新功能,如支持流/传感器数据,我们讨论了一些早期的科学用例驱动其设计。最后,我们提出了未来工作的路线图。
Modern computational research increasingly spans multiple, geographically distributed data centers and leverages instruments, experimental facilities and a network of national and regional cyberinfrastructure (CI). Tapis is an open-source API platform developed at the Texas Advanced Computing Center at the University of Texas at Austin to increase reproducibility and minimize time-to-solution for distributed computational experiments. Core features of Tapis include data management and code execution, a fine-grained permissions system enabling objects to be saved privately, shared with individuals or “published” to a community, and provenance endpoints exposing the detailed history Tapis collects on analyses, enabling workflows to be repeated and results reproduced. In this paper, we describe the evolution of the Tapis platform, from its origins in 2008, and discuss the growth and success of the project as well as challenges and limitations that have led to a new design effort, funded by the National Science Foundation in September of 2019. We present a detailed overview of the new system, including reference architecture and new features such as support for streaming/sensor data, and we discuss some of the early science use cases driving its design. We conclude with the roadmap for future work.