Gaudi components for concurrency: Concurrency for existing and future experiments

Gaudi components for concurrency: Concurrency for existing and future experiments
复制标题

用于并发的 Gaudi 组件:现有和未来实验的并发

DOI:
--
复制
发表时间:
2015
期刊:
影响因子:
--
通讯作者:
I. Shapoval
I. Shapoval
中科院分区:
--
文献类型:
--
作者:
M. Clemencic;Daniel Funke;B. Hegner;P. Mato;D. Piparo;I. Shapoval

文献摘要

被引文献

相似文献

HEP实验以不断增长的速度产生大量数据集。为了科普这些数据集带来的挑战,实验软件需要包含现代CPU提供的所有功能。随着内存/内核比率的降低,近年来的每个内核一个进程的方法变得不太可行。相反,需要利用具有细粒度并行性的多线程来从线程之间的内存共享中获益。高迪是一个独立于实验的数据处理框架,例如在欧洲核子研究中心的大型强子对撞机的ATLAS和LHCb实验中使用。它最初设计时只考虑顺序处理。在最近的努力中,框架已经扩展到允许多线程处理。这包括多个算法的并发调度组件-处理相同或多个事件,线程安全的数据存储访问和资源管理。在顺序的情况下,算法之间的关系被隐式地编码在它们预定的执行顺序中。对于并行处理,这些关系需要显式地表达,以便调度器能够在尊重算法之间的依赖性的同时利用最大并行性。因此,框架需要提供表达和自动跟踪这些依赖关系的方法。在本文中,我们提出了组件引入表达和跟踪算法的依赖关系,推导出一个优先约束的有向无环图,这是我们的算法复杂的调度方法的基础上,动态优先级的任务。我们引入了一个增量迁移路径,现有的实验对并行处理,并强调显式依赖的好处,即使在顺序的情况下,如健全检查和序列优化图分析。
HEP experiments produce enormous data sets at an ever-growing rate. To cope with the challenge posed by these data sets, experiments’ software needs to embrace all capabilities modern CPUs offer. With decreasing memory/core ratio, the one-process-per-core approach of recent years becomes less feasible. Instead, multi-threading with fine-grained parallelism needs to be exploited to benefit from memory sharing among threads. Gaudi is an experiment-independent data processing framework, used for instance by the ATLAS and LHCbexperiments at CERN's Large Hadron Collider. It has originally been designed with only sequential processing in mind. In a recent effort, the frame work has been extended to allow for multi-threaded processing. This includes components for concurrent scheduling of several algorithms - either processingthe same or multiple events, thread-safe data store access and resource management. In the sequential case, the relationships between algorithms are encoded implicitly in their pre-determined execution order. For parallel processing, these relationships need to be expressed explicitly, in order for the scheduler to be able to exploit maximum parallelism while respecting dependencies between algorithms. Therefore, means to express and automatically track these dependencies need to be provided by the framework. In this paper, we present components introduced to express and track dependencies of algorithms to deduce a precedence-constrained directed acyclic graph, which serves as basis for our algorithmically sophisticated scheduling approach for tasks with dynamic priorities. We introduce an incremental migration path for existing experiments towards parallel processing and highlight the benefits of explicit dependencies even in the sequential case, such as sanity checks and sequence optimization by graph analysis.