Protecting Synchronization Mechanisms of Parallel Big Data Kernels via Logging

Protecting Synchronization Mechanisms of Parallel Big Data Kernels via Logging
复制标题

DOI:
10.1109/tc.2021.3122993
复制
发表时间:
2022-09
影响因子:
3.7
通讯作者:
Travis LeCompte;Lu Peng;Xu Yuan;N. Tzeng
Travis LeCompte;Lu Peng;Xu Yuan;N. Tzeng
中科院分区:
计算机科学2区
文献类型:
--
作者:
Travis LeCompte;Lu Peng;Xu Yuan;N. Tzeng

文献摘要

相似文献

随着降低机器功耗的努力不断增加,容错能力变得更加令人担忧。这一点对于大规模计算尤其适用,在大规模计算中,由于软故障导致的执行失败浪费了大量的时间和资源。这些大规模应用程序本质上通常是并行的,并且依赖于专门为并行计算量身定做的控制结构,例如锁和障碍。虽然有很多关于弹性软件的研究,但据我们所知,没有一项研究侧重于保护这些并行控制结构。在这项工作中,我们提出了一种在并行应用中确保锁和栅栏的正确操作的方法。我们的方法跟踪并行部分中使用的内存位置,并检测对控制结构的违反。在检测到任何违规时,违规线程被回滚到结构的开头并重试,类似于事务内存系统中的回滚机制。我们在具有代表性的BigDataBitch内核样本上进行了测试,结果表明,对于基本的互斥锁和屏障,该方法的平均错误减少了93.6%,而在线程上的平均执行时间开销为6.55%。此外,我们提供了与事务存储方法的比较,并展示了高达57.5%的平均执行时间开销减少。
With the growing effort to reduce power consumption in machines, fault tolerance becomes more of a concern. This holds particularly for large-scale computing, where execution failures due to soft faults waste excessive time and resources. These large-scale applications are normally parallel in nature and rely on control structures tailored specifically for parallel computing, such as locks and barriers. While there are many studies on resilient software, to our knowledge none of them focus on protecting these parallel control structures. In this work, we present a method of ensuring the correct operation of both locks and barriers in parallel applications. Our method tracks the memory locations used within parallel sections and detects a violation of the control structures. Upon detecting any violation, the violating thread is rolled back to the beginning of the structure and reattempts it, similar to rollback mechanisms in transactional memory systems. We test the method on representative samples of the BigDataBench kernels and find it exhibits a mean error reduction of 93.6% for basic mutex locks and barriers with a mean 6.55% execution time overhead at 64 threads. Additionally, we provide a comparison to transactional memory methods and demonstrate up to a mean 57.5% execution time overhead reduction.