Harnessing Data Movement in Virtual Clusters for In-Situ Execution

Harnessing Data Movement in Virtual Clusters for In-Situ Execution
复制标题

利用虚拟集群中的数据移动进行原位执行

DOI:
--
复制
发表时间:
2019
影响因子:
5.3
通讯作者:
N. Podhorszki
N. Podhorszki
中科院分区:
计算机科学2区
文献类型:
--
作者:
Dan Huang;Qing Liu;S. Klasky;Jun Wang;J. Choi;Jeremy S. Logan;N. Podhorszki

文献摘要

被引文献

相似文献

由于数据量和速度的不断增加,百亿亿次级的大数据科学已转向原位范式,即大规模模拟与数据分析同时进行。通过原位范式,模拟产生的数据在内存中时就可进行处理,从而避免了缓慢的存储瓶颈。然而,如果不加以管理,在共享资源上同时运行模拟和分析可能会导致严重的争用,正如本文所展示的那样,这会导致模拟和分析的效率大幅降低。最近,诸如Linux容器等虚拟化技术已广泛应用于数据中心和物理集群,以便为包括科学模拟和数据分析在内的整合工作负载提供高效且灵活的资源配置。在本文中,我们研究如何在虚拟集群中为原位应用促进网络流量操控并减少网络上的相互干扰。为了在需要时动态分配网络带宽,我们采用基于季节性自回归移动平均模型(SARIMA)的技术来分析和预测模拟产生的消息传递接口(MPI)流量。尽管这可能是一种有效的技术,但单纯使用网络虚拟化可能会导致MPI作业内突发异步传输的性能下降。我们分析并解决了虚拟集群中的这种性能下降问题。
As a result of increasing data volume and velocity, Big Data science at exascale has shifted towards the in-situ paradigm, where large scale simulations run concurrently alongside data analytics. With in-situ, data generated from simulations can be processed while still in memory, thereby avoiding the slow storage bottleneck. However, running simulations and analytics together on shared resources will likely result in substantial contention if left unmanaged, as demonstrated in this work, leading to much reduced efficiency of simulations and analytics. Recently, virtualization technologies such as Linux containers have been widely applied to data centers and physical clusters to provide highly efficient and elastic resource provisioning for consolidated workloads including scientific simulations and data analytics. In this paper, we investigate to facilitate network traffic manipulation and reduce mutual interference on the network for in-situ applications in virtual clusters. In order to dynamically allocate the network bandwidth when it is needed, we adopt SARIMA-based techniques to analyze and predict MPI traffic issued from simulations. Although this can be an effective technique, the naïve usage of network virtualization can lead to performance degradation for bursty asynchronous transmissions within an MPI job. We analyze and resolve this performance degradation in virtual clusters.