Core Failure Mitigation in Integer Sum-of-Product Computations on Cloud Computing Systems

Core Failure Mitigation in Integer Sum-of-Product Computations on Cloud Computing Systems
复制标题

DOI:
10.1109/tmm.2016.2532603
复制
发表时间:
2016-02
影响因子:
7.3
通讯作者:
Ijeoma Anarado;Y. Andreopoulos
Ijeoma Anarado;Y. Andreopoulos
中科院分区:
计算机科学1区
文献类型:
--
作者:
Ijeoma Anarado;Y. Andreopoulos

文献摘要

相似文献

云计算系统中平均故障时间估计值的下降表明,在此类环境中运行的多媒体应用程序应该能够减轻运行时不断增加的核心故障数量。我们提出了一种新的前滚故障缓解方法,用于整数乘积和计算,重点是通用矩阵乘法(GEMM)和卷积/互相关(CONV)例程。我们的方法基于通过使用数字包装在输出的数字表示中产生冗余结果。这与需要一组单独的校验和(或重复)结果的所有现有前滚解决方案不同。与在 32 位整数表示上执行的整数乘积和实现相比,我们的建议将支持的最大输出位宽减少了 37.5%,这与多核故障缓解的校验和方法的位宽要求相当。在亚马逊 Web 服务弹性计算云 (AWS EC2) 的 c4.8xlarge 计算优化实例上运行最先进的 GEMM 和 CONV 例程进行的实验表明,所提出的方法能够减轻最多一个四核故障,同时实现处理吞吐量:1) 与传统的、容错的整数 GEMM 和 CONV 例程相当,2) 明显优于传统的、容错的整数 GEMM 和 CONV 例程。 基于校验和流的等效前滚故障缓解方法。此外,当在部署在 AWS EC2 Spot(即低成本但可终止)实例集群上的图像检索框架中使用时,我们的建议将导致:1)与基于校验和的等效方法相比,成本降低了 16%-23%;2)与 AWS EC2 按需上的传统容错处理相比,成本降低了 70% 以上(即成本较高,但成本较高)。 保证)实例。
The decreasing mean-time-to-failure estimates in cloud computing systems indicate that multimedia applications running on such environments should be able to mitigate an increasing number of core failures at runtime. We propose a new roll-forward failure-mitigation approach for integer sum-of-product computations, with emphasis on generic matrix multiplication (GEMM) and convolution/crosscorrelation (CONV) routines. Our approach is based on the production of redundant results within the numerical representation of the outputs via the use of numerical packing. This differs from all existing roll-forward solutions that require a separate set of checksum (or duplicate) results. Our proposal imposes 37.5% reduction in the maximum output bitwidth supported in comparison to integer sum-of-product realizations performed on 32-bit integer representations which is comparable to the bitwidth requirement of checksum-methods for multiple core failure mitigation. Experiments with state-of-the-art GEMM and CONV routines running on a c4.8xlarge compute-optimized instance of amazon web services elastic compute cloud (AWS EC2) demonstrate that the proposed approach is able to mitigate up to one quadcore failure while achieving processing throughput that is: 1) comparable to that of the conventional, failure-intolerant, integer GEMM and CONV routines, 2) substantially superior to that of the equivalent roll-forward failure-mitigation method based on checksum streams. Furthermore, when used within an image retrieval framework deployed over a cluster of AWS EC2 spot (i.e., low-cost albeit terminatable) instances, our proposal leads to: 1) 16%-23% cost reduction against the equivalent checksum-based method and 2) more than 70% cost reduction against conventional failure-intolerant processing on AWS EC2 on-demand (i.e., higher-cost albeit guaranteed) instances.