The Quest to solve the HL-LHC data access puzzle

The Quest to solve the HL-LHC data access puzzle
复制标题

解决 HL-LHC 数据访问难题的探索

DOI:
--
复制
发表时间:
2020
影响因子:
--
通讯作者:
F. Wuerthwein
F. Wuerthwein
中科院分区:
--
文献类型:
--
作者:
X. Espinal;S. Jézéquel;M. Schulz;A. Sciabà;I. Vukotic;F. Wuerthwein

文献摘要

被引文献

相似文献

HL-LHC将面临WLCG社区巨大的数据存储、管理和访问挑战。这些既经济又技术。在WLCG-DOMA访问工作组中,实验成员和站点管理人员考虑到我们社区给出的边界条件,探索了不同的数据访问和存储策略模型,以降低成本和复杂性。已经对其中几个方案进行了定量评估,例如数据湖模型和当前计算模型在资源需求、成本和操作复杂性方面的增量改进。为了更好地深入理解这些模型,对当前数据访问的轨迹进行了分析,并对新概念的影响进行了模拟。同时,对所需的技术进行了评价。这些都是在小型和大型的测试平台和生产环境中完成的。我们将概述工作组的活动和结果,描述模型并总结技术评估的结果,重点关注以数据湖的形式进行存储整合的影响,其中使用流缓存已成为减少延迟和带宽限制影响的成功方法。我们将描述这些方法在不同环境和使用场景中的经验和评估。此外,我们将介绍基于实验数据访问痕迹的分析和建模工作的结果。
HL-LHC will confront the WLCG community with enormous data storage, management and access challenges. These are as much technical as economical. In the WLCG-DOMA Access working group, members of the experiments and site managers have explored different models for data access and storage strategies to reduce cost and complexity, taking into account the boundary conditions given by our community.Several of these scenarios have been evaluated quantitatively, such as the Data Lake model and incremental improvements of the current computing model with respect to resource needs, costs and operational complexity.To better understand these models in depth, analysis of traces of current data accesses and simulations of the impact of new concepts have been carried out. In parallel, evaluations of the required technologies took place. These were done in testbed and production environments at small and large scale.We will give an overview of the activities and results of the working group, describe the models and summarise the results of the technology evaluation focusing on the impact of storage consolidation in the form of Data Lakes, where the use of streaming caches has emerged as a successful approach to reduce the impact of latency and bandwidth limitation.We will describe the experience and evaluation of these approaches in different environments and usage scenarios. In addition we will present the results of the analysis and modelling efforts based on data access traces of the experiments.