EOS architectural evolution and strategic development directions

EOS architectural evolution and strategic development directions
复制标题

EOS架构演进及战略发展方向

DOI:
10.1051/epjconf/202024504009
复制
发表时间:
2020
影响因子:
--
通讯作者:
Elvin A. Sindrilaru
Elvin A. Sindrilaru
中科院分区:
--
文献类型:
--
作者:
G. Bitzes;F. Luchetti;A. Manzi;M. Patrascoiu;A. Peters;M. Simon;Elvin A. Sindrilaru

文献摘要

被引文献

相似文献

EOS[1]是CERN的主要存储系统,为物理实验和CERN基础设施的常规用户提供了数百PB的容量。自2010年首次部署以来,EOS已经发展并适应了不断增长的存储容量要求、用户友好的POSIX式交互体验以及协作应用程序等新模式以及同步和共享功能带来的挑战。 在软件堆栈的不同级别上克服这些挑战意味着要为命名空间子系统提出新的体系结构,完全重新设计EOS熔丝模块,并调整其余组件,如排水、LRU引擎、文件系统一致性检查等,以确保稳定和可预测的性能。在本文中,我们详细介绍了触发所有这些更改的问题以及我们所做的软件设计选择。 在白皮书的最后部分,我们将重点放在需要立即改进的领域,以确保最终用户的无缝体验以及提高服务的整体可用性。其中一些更改影响深远,旨在简化部署模型,更重要的是简化在管理数千个磁盘的系统中处理(非/)瞬时错误时的操作负载。
EOS [1] is the main storage system at CERN providing hundreds of PB of capacity to both physics experiments and also regular users of the CERN infrastructure. Since its first deployment in 2010, EOS has evolved and adapted to the challenges posed by ever-increasing requirements for storage capacity, user-friendly POSIX-like interactive experience and new paradigms like collaborative applications along with sync and share capabilities. Overcoming these challenges at various levels of the software stack meant coming up with a new architecture for the namespace subsystem, completely redesigning the EOS FUSE module and adapting the rest of the components like draining, LRU engine, file system consistency check and others, to ensure a stable and predictable performance. In this paper we detail the issues that triggered all these changes along with the software design choices that we made. In the last part of the paper, we move our focus to the areas that need immediate improvements in order to ensure a seamless experience for the end-user along with increased over-all availability of the service. Some of these changes have far-reaching effects and are aimed at simplifying both the deployment model but more importantly the operational load when dealing with (non/)transient errors in a system managing thousands of disks.