Block-parallel data analysis with DIY2

Block-parallel data analysis with DIY2
复制标题

DOI:
10.1109/ldav.2016.7874307
复制
发表时间:
2016-05
期刊:
2016 IEEE 6th Symposium on Large Data Analysis and Visualization (LDAV)
影响因子:
--
通讯作者:
D. Morozov;T. Peterka
D. Morozov;T. Peterka
中科院分区:
其他
文献类型:
--
作者:
D. Morozov;T. Peterka

文献摘要

被引文献

相似文献

DIY 2是一个编程模型和运行时,用于分布式内存机器上的块并行分析。它的主要抽象是块结构的数据并行性:数据被分解为块;块被分配给处理元素(进程或线程);计算被描述为在这些块上的迭代,块之间的通信由可重用模式定义。通过以这种一般形式表示计算,DIY 2运行时可以自由地优化块在慢速和快速存储器(磁盘和闪存与DRAM)之间的移动,并使用多个线程并发执行驻留在内存中的块。这使得相同的程序能够执行核内、核外、串行、并行、单线程、多线程或其组合。本文介绍了DIY 2编程模型的主要功能的实现和优化,以提高性能。在完整的分析代码上评估DIY 2。
DIY2 is a programming model and runtime for block-parallel analytics on distributed-memory machines. Its main abstraction is block-structured data parallelism: data are decomposed into blocks; blocks are assigned to processing elements (processes or threads); computation is described as iterations over these blocks, and communication between blocks is defined by reusable patterns. By expressing computation in this general form, the DIY2 runtime is free to optimize the movement of blocks between slow and fast memories (disk and flash vs. DRAM) and to concurrently execute blocks residing in memory with multiple threads. This enables the same program to execute in-core, out-of-core, serial, parallel, single-threaded, multithreaded, or combinations thereof. This paper describes the implementation of the main features of the DIY2 programming model and optimizations to improve performance. DIY2 is evaluated on complete analysis codes.