Enabling scalability-sensitive speculative parallelization for FSM computations

Enabling scalability-sensitive speculative parallelization for FSM computations
复制标题

为 FSM 计算启用可扩展性敏感的推测并行化

DOI:
10.1145/3079079.3079082
复制
发表时间:
2017
期刊:
Proceedings of the International Conference on Supercomputing
影响因子:
--
通讯作者:
S. Song
S. Song
中科院分区:
--
文献类型:
--
作者:
Junqiao Qiu;Zhijia Zhao;Bo Wu;Abhinav Vishnu;S. Song

文献摘要

被引文献

相似文献

有限状态机(FSM)是许多应用程序的骨干,但由于其固有的依赖性,很难并行化。投机性FSM并行化在多功能机器上表现出了有望,最多八个核心。但是,随着硬件并行性的增长(例如,Xeon Phi具有多达288个逻辑核心),一个基本问题提出了:投机性FSM平行量表如何随着核心数量的增加而产生?在没有回答这个问题的情况下,现有的投机FSM并行化方法只是选择使用所有可用的内核,这可能不仅浪费了计算资源,而且会导致次优性能。在这项工作中,我们对投机FSM并行化进行了系统的可伸缩性分析。与许多其他可以通过经典Amdahl定律或其简单扩展模型的并行化不同,由于投机性的非确定性和拼写错误的成本变化,投机FSM平行化尺度在不常规上尺度。为了应对这些挑战,这项工作引入了一系列可扩展性模型,这些模型是针对特定FSM和基础体系结构的属性进行定制的。这些模型首次精确地捕获了针对不同FSM计算的投机并行化的可扩展性,并清楚地表明了通过投机FSM并行化以实现最佳性能的核心数量的“甜点”。为了使可伸缩性模型实用,我们开发了S3,这是用于FSM并行化的对可伸缩性敏感的投机框架。对于任何给定的FSM,S3可以自动表征其属性并分析其可扩展性,因此指导投机性并行化,以实现最佳性能和更有效地利用计算资源。对不同FSM和体系结构的评估证实了所提出的模型的准确性,并表明S3实现了明显的加速(最高5倍)和节能(最高77%)。
Finite state machines (FSMs) are the backbone of many applications, but are difficult to parallelize due to their inherent dependencies. Speculative FSM parallelization has shown promise on multicore machines with up to eight cores. However, as hardware parallelism grows (e.g., Xeon Phi has up to 288 logical cores), a fundamental question raises: How does the speculative FSM parallelization scale as the number of cores increases? Without answering this question, existing methods for speculative FSM parallelization simply choose to use all available cores, which might not only waste computing resources, but also result in suboptimal performance. In this work, we conduct a systematic scalability analysis for speculative FSM parallelization. Unlike many other parallelizations which can be modeled by the classic Amdahl's law or its simple extensions, speculative FSM parallelization scales unconventionally due to the non-deterministic nature of speculation and the cost variations of misspeculation. To address these challenges, this work introduces a spectrum of scalability models that are customized to the properties of specific FSMs and the underlying architecture. The models, for the first time, precisely capture the scalability of speculative parallelization for different FSM computations, and clearly show the existence of a "sweet spot" in terms of the number of cores employed by the speculative FSM parallelization to achieve the optimal performance. To make the scalability models practical, we develop S3, a scalability-sensitive speculation framework for FSM parallelization. For any given FSM, S3 can automatically characterize its properties and analyze its scalability, hence guide speculative parallelization towards the optimal performance and more efficient use of computing resources. Evaluations on different FSMs and architectures confirm the accuracy of the proposed models and show that S3 achieves significant speedup (up to 5X) and energy savings (up to 77%).