Skew-resistant parallel processing of feature-extracting scientific user-defined functions

Skew-resistant parallel processing of feature-extracting scientific user-defined functions
复制标题

DOI:
10.1145/1807128.1807140
复制
发表时间:
2010-06
期刊:
--
影响因子:
--
通讯作者:
YongChul Kwon;M. Balazinska;Bill Howe;J. Rolia
YongChul Kwon;M. Balazinska;Bill Howe;J. Rolia
中科院分区:
其他
文献类型:
--
作者:
YongChul Kwon;M. Balazinska;Bill Howe;J. Rolia

文献摘要

被引文献

相似文献

如今,科学家有能力以前所未有的规模和速度生成数据,因此,他们必须越来越多地转向并行数据处理引擎来执行分析。然而,这些引擎的简单执行模型可能导致难以实现有效的科学分析算法。特别是,许多科学分析需要从表示为多维数组或多维空间中的点的数据中提取特征。这些应用程序表现出显着的计算偏差,其中不同分区的运行时间不仅仅取决于输入大小,因此可能会发生巨大且不可预测的变化。在本文中,我们提出了 SkewReduce,这是一个在 Hadoop 之上实现的新系统,使用户能够轻松表达特征提取分析并高效执行它们。 SkewReduce 系统的核心是一个优化器,由用户定义的成本函数参数化,它确定如何最好地划分输入数据以最小化计算偏差。对来自两个不同科学领域的真实数据进行的实验表明,与简单的实现相比,我们的方法可以将执行时间提高多达 8 倍。
Scientists today have the ability to generate data at an unprecedented scale and rate and, as a result, they must increasingly turn to parallel data processing engines to perform their analyses. However, the simple execution model of these engines can make it difficult to implement efficient algorithms for scientific analytics. In particular, many scientific analytics require the extraction of features from data represented as either a multidimensional array or points in a multidimensional space. These applications exhibit significant computational skew, where the runtime of different partitions depends on more than just input size and can therefore vary dramatically and unpredictably. In this paper, we present SkewReduce, a new system implemented on top of Hadoop that enables users to easily express feature extraction analyses and execute them efficiently. At the heart of the SkewReduce system is an optimizer, parameterized by user-defined cost functions, that determines how best to partition the input data to minimize computational skew. Experiments on real data from two different science domains demonstrate that our approach can improve execution times by a factor of up to 8 compared to a naive implementation.