PyCOMPSs: Parallel computational workflows in Python

PyCOMPSs: Parallel computational workflows in Python
复制标题

PyCOMPS:Python 中的并行计算工作流程

DOI:
--
复制
发表时间:
2016
期刊:
The international journal of high performance computing applications
影响因子:
--
通讯作者:
Jesús Labarta
Jesús Labarta
中科院分区:
--
文献类型:
--
作者:
E. Tejedor;Y. Becerra;Guillem Alomar;A. Queralt;Rosa M. Badia;J. Torres;Toni Cortes;Jesús Labarta

文献摘要

被引文献

相似文献

在过去的几年里,Python编程语言在科学计算中的使用一直在增长。事实上,它是紧凑和可读的,其完整的科学图书馆是两个重要的特点,有利于它的采用。尽管如此,Python仍然缺乏在分布式基础设施上轻松并行化通用脚本的解决方案,因为目前的替代方案大多需要使用API进行消息传递或仅限于并行计算。在这个意义上,本文介绍了PyCOMPSs,这是一个促进Python中并行计算工作流开发的框架。在这种方法中,用户以顺序的方式编写脚本,并装饰要作为异步并行任务运行的函数。运行时系统负责利用脚本固有的并发性,检测任务之间的数据依赖关系,并将它们生成到可用资源。此外,我们展示了如何在大数据存储架构之上构建这种编程模型,其中存储在后端的数据以持久对象的形式从应用程序中抽象和访问。
The use of the Python programming language for scientific computing has been gaining momentum in the last years. The fact that it is compact and readable and its complete set of scientific libraries are two important characteristics that favour its adoption. Nevertheless, Python still lacks a solution for easily parallelizing generic scripts on distributed infrastructures, since the current alternatives mostly require the use of APIs for message passing or are restricted to embarrassingly parallel computations. In that sense, this paper presents PyCOMPSs, a framework that facilitates the development of parallel computational workflows in Python. In this approach, the user programs her script in a sequential fashion and decorates the functions to be run as asynchronous parallel tasks. A runtime system is in charge of exploiting the inherent concurrency of the script, detecting the data dependencies between tasks and spawning them to the available resources. Furthermore, we show how this programming model can be built on top of a Big Data storage architecture, where the data stored in the backend is abstracted and accessed from the application in the form of persistent objects.