ATLAS BigPanDA monitoring

ATLAS BigPanDA monitoring
复制标题

ATLAS BigPanDA 监控

DOI:
--
复制
发表时间:
2018
期刊:
Journal of Physics: Conference Series
影响因子:
--
通讯作者:
T. Wenaus
T. Wenaus
中科院分区:
--
文献类型:
--
作者:
A. Alekseev;A. Klimentov;T. Korchuganova;S. Padolski;T. Wenaus

文献摘要

被引文献

相似文献

BigPanDA监控是一个Web应用程序,它提供对生产和分布式分析(PANDA)系统对象状态的各种处理和表示。通过分析数以亿计的计算实体,如事件或作业,BigPanDA监控以实时模式构建不同规模和级别的抽象报告。提供的信息使用户能够深入了解具体事件失败的原因,或观察总体情况,如跟踪计算核心和卫星性能或整个生产活动的进展。熊猫系统最初是为ATLAS实验而开发的。目前,它每天管理分布在全球170个计算中心的200多万个作业的执行。BigPanDA是其于2014年年中投入使用的核心组件,现在是ATLAS用户关于其计算状态的主要信息来源,也是轮班人员、操作员和管理人员的决策支持信息来源。在这项工作中,我们描述了BigPanDA监测体系结构的演变、现状和发展计划。
BigPanDA monitoring is a web application that provides various processing and representation of the Production and Distributed Analysis (PanDA) system objects states. Analysing hundreds of millions of computation entities, such as an event or a job, BigPanDA monitoring builds different scales and levels of abstraction reports in real time mode. Provided information allows users to drill down into the reason of a concrete event failure or observe the broad picture such as tracking the computation nucleus and satellites performance or the progress of a whole production campaign. PanDA system was originally developed for the ATLAS experiment. Currently, it manages execution of more than 2 million jobs distributed over 170 computing centers worldwide on daily basis. BigPanDA is its core component commissioned in the middle of 2014 and now is the primary source of information for ATLAS users about the state of their computations and the source of decision support information for shifters, operators and managers. In this work, we describe the evolution of the architecture, current status and plans for the development of the BigPanDA monitoring.