Managing the ATLAS Grid through Harvester

Managing the ATLAS Grid through Harvester
复制标题

通过 Harvester 管理 ATLAS 网格

DOI:
--
复制
发表时间:
2020
影响因子:
--
通讯作者:
Nicolò Magini
Nicolò Magini
中科院分区:
--
文献类型:
--
作者:
F. B. Megino;A. Alekseev;F. Berghaus;D. Cameron;K. De;A. Filipčič;I. Glushkov;Fahui Lin;T. Maeno;Nicolò Magini

文献摘要

被引文献

相似文献

阿特拉斯计算管理公司已将所有计算资源迁移到Panda的新工作负载提交引擎Harvester,作为LHC Run 3和LHC Run 4的关键里程碑。这一贡献将集中在网格向Harvester的迁移上。我们基于CERN IT的常见产品(例如OpenStack虚拟机和按需数据库)构建了一个冗余架构,以运行必要的收割机和HTCondor服务,能够承受电网上每天O(100万)工作人员的负载。我们已经逐个地区审查了ATLAS网格,并尽可能地避免盲目提交工作进程,即多个队列(例如,单核、多核、高内存)竞争站点上的资源。相反,我们已经转向更智能的模型,这些模型使用来自中央熊猫工作负载管理系统的信息和优先级,并将每个类别的适当数量的工作人员流到统一队列中,同时保持对作业的后期绑定。我们还将介绍我们增强的监测和分析框架。员工和作业信息与CERN IT部门提供的ElasticSearch存储库的同步延迟最小,我们可以在其中与仪表板交互,以跟踪提交进度、发现站点问题(例如,损坏的计算元素)或发现空的员工。其结果是通过智能、内置的资源监控,更有效地使用网格资源。
ATLAS Computing Management has identified the migration of all computing resources to Harvester, PanDA’s new workload submission engine, as a critical milestone for LHC Run 3 and 4. This contribution will focus on the Grid migration to Harvester. We have built a redundant architecture based on CERN IT’s common offerings (e.g. Openstack Virtual Machines and Database on Demand) to run the necessary Harvester and HTCondor services, capable of sustaining the load of O(1M) workers on the Grid per day. We have reviewed the ATLAS Grid region by region and moved as much possible away from blind worker submission, where multiple queues (e.g. single core, multi core, high memory) compete for resources on a site. Instead we have migrated towards more intelligent models that use information and priorities from the central PanDA workload management system and stream the right number of workers of each category to a unified queue while keeping late binding to the jobs. We will also describe our enhanced monitoring and analytics framework. Worker and job information is synchronized with minimal delays to a CERN IT provided ElasticSearch repository, where we can interact with dashboards to follow submission progress, discover site issues (e.g. broken Compute Elements) or spot empty workers. The result is a much more efficient usage of the Grid resources with smart, built-in monitoring of resources.