Multi-core Implementations of the Concurrent Collections Programming Model

Multi-core Implementations of the Concurrent Collections Programming Model
复制标题

并发集合编程模型的多核实现

DOI:
--
复制
发表时间:
2008
期刊:
影响因子:
--
通讯作者:
Leo Treggiari
Leo Treggiari
中科院分区:
--
文献类型:
--
作者:
Zoran Budimlic;Aparna Chandramowlishwaran;K. Knobe;Geoff N. Lowney;Vivek Sarkar;Leo Treggiari

文献摘要

被引文献

相似文献

在本文中,我们介绍了并发收集编程模型,它建立在TStreams [8]的过去工作的基础上。在这个模型中,程序是按照高级应用程序特定的操作编写的。这些操作仅根据其语义约束进行部分排序。这些偏序对应于数据流和控制流。这种方法支持重要的关注点分离。实现并行程序涉及两个角色。一个是领域专家的角色,开发人员的兴趣和专业知识是在应用领域,如金融,基因组学或数值分析。另一个是调优专家,他们的兴趣和专长是性能,包括特定平台上的性能。这些可能是不同的个体,也可能是处于应用程序开发不同阶段的同一个体。调优专家实际上可以是软件(例如静态或动态优化编译器)。并发集合编程模型将领域专家的工作(计算语义的表达)与调优专家的工作(实际并行性到特定体系结构的选择和映射)分开。这种分离简化了领域专家的任务。用这种语言编写不需要任何关于并行性的推理或对目标体系结构的理解。领域专家只关心他或她的专业领域(应用程序的语义)。这种分离还简化了调优专家的工作。调优专家被给予最大可能的自由来将计算映射到目标体系结构上,并且不需要对域有任何理解(编译器通常是这种情况)。我们描述了并发收集编程模型的两个实现。一个是基于英特尔®线程构建块的英特尔® C/C++并发集合。另一个是Rice大学Habanero项目的基于X10的实现。我们比较的实现,显示多核SMP机器上执行相同的并发收集应用程序时,Cholesky因式分解,在这两种方法实现的结果。
In this paper we introduce the Concurrent Collections programming model, which builds on past work on TStreams [8]. In this model, programs are written in terms of high-level application-specific operations. These operations are partially ordered according to only their semantic constraints. These partial orderings correspond to data flow and control flow. This approach supports an important separation of concerns. There are two roles involved in implementing a parallel program. One is the role of a domain expert, the developer whose interest and expertise is in the application domain, such as finance, genomics, or numerical analysis. The other is the tuning expert, whose interest and expertise is in performance, including performance on a particular platform. These may be distinct individuals or the same individual at different stages in application development. The tuning expert may in fact be software (such as a static or dynamic optimizing compiler). The Concurrent Collections programming model separates the work of the domain expert (the expression of the semantics of the computation) from the work of the tuning expert (selection and mapping of actual parallelism to a specific architecture). This separation simplifies the task of the domain expert. Writing in this language does not require any reasoning about parallelism or any understanding of the target architecture. The domain expert is concerned only with his or her area of expertise (the semantics of the application). This separation also simplifies the work of the tuning expert. The tuning expert is given the maximum possible freedom to map the computation onto the target architecture and is not required to have any understanding of the domain (as is often the case for compilers). We describe two implementations of the Concurrent Collections programming model. One is IntelR © Concurrent Collections for C/C++ based on IntelR © Threaded Building Blocks. The other is an X10-based implementation from the Habanero project at Rice University. We compare the implementations by showing the results achieved on multi-core SMP machines when executing the same Concurrent Collections application, Cholesky factorization, in both these approaches.