Pilot factory – a Condor-based system for scalable Pilot Job generation in the Panda WMS framework

Pilot factory – a Condor-based system for scalable Pilot Job generation in the Panda WMS framework
复制标题

试点工厂——基于 Condor 的系统,用于在 Panda WMS 框架中生成可扩展的试点作业

DOI:
10.1088/1742-6596/219/6/062041
复制
发表时间:
2010
期刊:
Journal of Physics: Conference Series
影响因子:
--
通讯作者:
M. Potekhin
M. Potekhin
中科院分区:
--
文献类型:
--
作者:
Po;M. Potekhin

文献摘要

被引文献

相似文献

Panda 工作负载管理系统是围绕 Pilot Job 的概念设计的 - 一个有效负载可执行文件的“智能包装器”,可以在从服务器拉取有效负载并执行之前探测远程工作节点上的环境。这种设计可以改进日志记录和监控功能以及工作负载管理的灵活性。在网格环境(例如开放科学网格)中,Panda Pilot 作业通过最终依赖于 Condor-G 的机制提交到远程站点。我们的经验表明,在大量 Panda 作业同时路由到特定远程站点的情况下,由于 Pilot 作业提交而导致集群头节点负载增加,可能会导致整体缺乏可扩展性。我们针对这个问题开发了一个受 Condor 启发的解决方案,它使用基于 schedd 的 glidein,其任务是将飞行员重定向到本机批处理系统。一旦安装并运行了 glidein schedd,它就可以像本地 schedd 一样使用,因此,从用户的角度来看,提交的 Pilot 与提交到本地 Condor 池的作业非常相似。
The Panda Workload Management System is designed around the concept of the Pilot Job – a "smart wrapper" for the payload executable that can probe the environment on the remote worker node before pulling down the payload from the server and executing it. Such design allows for improved logging and monitoring capabilities as well as flexibility in Workload Management. In the Grid environment (such as the Open Science Grid), Panda Pilot Jobs are submitted to remote sites via mechanisms that ultimately rely on Condor-G. As our experience has shown, in cases where a large number of Panda jobs are simultaneously routed to a particular remote site, the increased load on the head node of the cluster, which is caused by the Pilot Job submission, may lead to overall lack of scalability. We have developed a Condor-inspired solution to this problem, which is using the schedd-based glidein, whose mission is to redirect pilots to the native batch system. Once a glidein schedd is installed and running, it can be utilized exactly the same way as local schedds and therefore, from the user's perspective, Pilots thus submitted are quite similar to jobs submitted to the local Condor pool.