ATLAS Data Carousel

ATLAS Data Carousel
复制标题

ATLAS 数据轮播

DOI:
--
复制
发表时间:
2020
影响因子:
--
通讯作者:
Xin Zhao
Xin Zhao
中科院分区:
--
文献类型:
--
作者:
M. Barisits;M. Borodin;A. Girolamo;J. Elmsheuser;D. Golubkov;A. Klimentov;M. Lassnig;T. Maeno;R. Walker;Xin Zhao

文献摘要

被引文献

相似文献

欧洲核子研究中心大型强子对撞机的ATLAS实验以原始和派生数据格式存储了全球150多个网格站点的探测器和模拟数据,目前磁盘上的总容量约为200PB,磁带上的总容量约为250PB。由于不同的计算工作流,数据具有不同的访问特征,并且可以从不同的介质访问,例如远程I/O、硬盘驱动器上的磁盘缓存或SSD。此外,较大的数据中心通过磁带系统提供大部分离线存储功能。对于高亮度大型强子对撞机(HL-LHC),在预算持平的假设下,估计的数据存储需求比目前预测的可用资源大几倍。在计算方面,ATLAS分布式计算在过去几年中凭借高性能和高吞吐量的计算集成以及在蒙特卡洛模拟中使用机会计算资源而非常成功。另一方面,对等的机会主义存储并不存在。Atlas启动了Data Carousel项目,以增加更便宜的存储(即磁带甚至商业存储)的使用率,因此它并不局限于磁带技术。Data Carousel协调工作负载管理、数据管理和存储服务之间的数据处理,并将大量数据驻留在离线存储上。该处理是通过将输入的滑动窗口分级并迅速处理到更快的缓冲存储器上来执行的,从而使得在任何时刻都只有一小部分输入数据可用。通过这个项目,我们的目标是证明这是显著降低存储成本的自然方式。该项目的第一阶段于2018年秋季开始,与网站存档系统的I/O测试有关。第二阶段现在需要将工作量和数据管理系统紧密结合起来。此外,Data Carousel还研究从磁带运行多个计算工作流的可行性。该项目进展非常顺利,本文中提出的结果将在大型强子对撞机第三次运行之前使用。
The ATLAS experiment at CERN’s LHC stores detector and simulation data in raw and derived data formats across more than 150 Grid sites world-wide, currently in total about 200PB on disk and 250PB on tape. Data have different access characteristics due to various computational workflows, and can be accessed from different media, such as remote I/O, disk cache on hard disk drives or SSDs. Also, larger data centers provide the majority of offline storage capability via tape systems. For the HighLuminosity LHC (HL-LHC), the estimated data storage requirements are several factors bigger than the present forecast of available resources, based on a flat budget assumption. On the computing side, ATLAS Distributed Computing was very successful in the last years with high performance and high throughput computing integration and in using opportunistic computing resources for the Monte Carlo simulation. On the other hand, equivalent opportunistic storage does not exist. ATLAS started the Data Carousel project to increase the usage of less expensive storage, i.e. tapes or even commercial storage, so it is not limited to tape technologies exclusively. Data Carousel orchestrates data processing between workload management, data management, and storage services with the bulk data resident on offline storage. The processing is executed by staging and promptly processing a sliding window of inputs onto faster buffer storage, such that only a small percentage of input data are available at any one time. With this project, we aim to demonstrate that this is the natural way to dramatically reduce our storage cost. The first phase of the project was started in the fall of 2018 and was related to I/O tests of the sites archiving systems. Phase II now requires a tight integration of the workload and data management systems. Additionally, the Data Carousel studies the feasibility to run multiple computing workflows from tape. The project is progressing very well and the results presented in this document will be used before the LHC Run 3.