Towards a NoOps Model for WLCG

Towards a NoOps Model for WLCG
复制标题

迈向 WLCG 的 NoOps 模型

DOI:
10.1051/epjconf/202024507024
复制
发表时间:
2020
期刊:
24th International Conference on Computing in High Energy and Nuclear Physics (CHEP 2019
影响因子:
--
通讯作者:
Wu, Wenjing
Wu, Wenjing
中科院分区:
--
文献类型:
--
作者:
Gardner, Robert;Bryant, Lincoln;Stephen, Judith;Vukotic, Ilija;Weaver, Christopher;Wu, Wenjing

文献摘要

参考文献

相似文献

提供全球计算基础设施(如WLCG)的最昂贵因素之一是部署、集成和操作支持协作计算、数据共享和交付以及极端规模数据集分析的分布式服务的人力。此外,推出全球软件更新、引入新的服务组件或原型化需要跨多个设施协调部署的新系统所需的时间通常会因通信延迟、工作人员可用性以及在许多情况下定制服务操作所需的专业知识而增加。虽然WLCG(以及在整个HEP中实现的分布式系统)是一个全球服务平台,但它缺乏现代平台即服务的能力和灵活性,包括持续集成/持续交付(CI/CD)方法,开发运营能力(DevOps,其中开发人员在实际生产基础设施中承担更直接的角色)和自动化。最重要的是,减少所需培训、定制服务专业知识和整个基础设施(尤其是资源端点(站点))的操作工作的工具在当前模型中完全不存在。在本文中,我们将探讨潜在的NoOpmodels在这种情况下的想法和问题:什么是现实的组织政策和约束?如何在团队和设施之间组织运营责任?技术差距是什么?社会和网络安全挑战是什么?相反,NoOps模型为创新和加速HL-LHC时代所需的新服务的交付速度带来了什么优势?我们将在提供支持IRIS-HEP DOMA R&D的数据传送网络的背景下描述沿着沿着这些路线的初始工作。
One of the most costly factors in providing a global computing infrastructure such as the WLCG is the human effort in deployment, integration, and operation of the distributed services supporting collaborative computing, data sharing and delivery, and analysis of extreme scale datasets. Furthermore, the time required to roll out global software updates, introduce new service components, or prototype novel systems requiring coordinated deployments across multiple facilities is often increased by communication latencies, staff availability, and in many cases expertise required for operations of bespoke services. While the WLCG (and distributed systems implemented throughout HEP) is a global service platform, it lacks the capability and flexibility of a modern platform-as-a-service including continuous integration/continuous delivery (CI/CD) methods, development-operations capabilities (DevOps, where developers assume a more direct role in the actual production infrastructure), and automation. Most importantly, tooling which reduces required training, bespoke service expertise, and the operational effort throughout the infrastructure, most notably at the resource endpoints (sites), is entirely absent in the current model. In this paper, we explore ideas and questions around potentialNoOpsmodels in this context: what is realistic given organizational policies and constraints? How should operational responsibility be organized across teams and facilities? What are the technical gaps? What are the social and cybersecurity challenges? Conversely what advantages does a NoOps model deliver for innovation and for accelerating the pace of delivery of new services needed for the HL-LHC era? We will describe initial work along these lines in the context of providing a data delivery network supporting IRIS-HEP DOMA R&D.
DOI: 10.1051/epjconf/201921403010
发表时间: 2019
影响因子: --
作者:
J. Elmsheuser;A. Girolamo
通讯作者: A. Girolamo
DOI: 10.7717/peerj-cs.144
发表时间: 2018
期刊: PeerJ. Computer science
影响因子: --
作者:
Chard K;Dart E;Foster I;Shifflett D;Tuecke S;Williams J
通讯作者: Williams J
DOI: 10.1051/epjconf/201921403049
发表时间: 2019
影响因子: --
作者:
Fernando Barreiro Magino;D. Cameron;Alessandro di Girolamo;A. Filipčič;I. Glushkov;F. Legger;T. Maeno;R. Walker
通讯作者: R. Walker
构建 SLATE 平台
DOI: 10.1145/3219104.3219144
发表时间: 2018
期刊: Proceedings of the Practice and Experience on Advanced Research Computing
影响因子: --
作者:
Breen, Joe;McKee, Shawn;Riedel, Benedikt;Stidd, Jason;Truong, Luan;Vukotic, Ilija;Bryant, Lincoln;Carcassi, Gabriele;Chen, Jiahui;Gardner, Robert W.
通讯作者: Gardner, Robert W.