I/O load balancing for big data HPC applications

I/O load balancing for big data HPC applications
复制标题

DOI:
10.1109/bigdata.2017.8257931
复制
发表时间:
2017-12
期刊:
2017 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
A. Paul;Anuj Kumar Goyal;Feiyi Wang;S. Oral;A. Butt;Michael J. Brim;Sangeetha B. Srinivasa
A. Paul;Anuj Kumar Goyal;Feiyi Wang;S. Oral;A. Butt;Michael J. Brim;Sangeetha B. Srinivasa
中科院分区:
其他
文献类型:
--
作者:
A. Paul;Anuj Kumar Goyal;Feiyi Wang;S. Oral;A. Butt;Michael J. Brim;Sangeetha B. Srinivasa

文献摘要

被引文献

相似文献

高性能计算(High Performance Computing, HPC)大数据问题需要高效的分布式存储系统。然而,在规模上,这样的存储系统经常遇到负载不平衡和资源争用,这是由于两个因素:科学应用程序I/O的突然性;以及没有集中仲裁和控制的复杂I/O路径。例如,现有的Lustre并行文件系统支持许多HPC中心,它包含通过自定义网络拓扑连接的许多组件,并满足大量用户和应用程序的不同需求。因此,一些存储服务器可能比其他存储服务器负载更大,从而产生瓶颈并降低整体应用程序I/O性能。现有的解决方案通常侧重于每个应用程序的负载平衡,因此在缺乏系统全局视图的情况下效率不高。在本文中,我们提出了一种数据驱动的方法来大规模地平衡I/O服务器的负载,目标是Lustre部署。为此,我们在Lustre Metadata Server上设计了一个全局映射器,它从I/O路径上的关键存储组件收集运行时统计数据,并应用马尔可夫链建模和最小成本最大流算法来确定数据应该放置在哪里。使用真实系统模拟器和真实设置进行的评估表明,我们的方法可以产生更好的负载平衡,从而可以提高端到端性能。
High Performance Computing (HPC) big data problems require efficient distributed storage systems. However, at scale, such storage systems often experience load imbalance and resource contention due to two factors: the bursty nature of scientific application I/O; and the complex I/O path that is without centralized arbitration and control. For example, the extant Lustre parallel file system-that supports many HPC centers-comprises numerous components connected via custom network topologies, and serves varying demands of a large number of users and applications. Consequently, some storage servers can be more loaded than others, which creates bottlenecks and reduces overall application I/O performance. Existing solutions typically focus on per application load balancing, and thus are not as effective given their lack of a global view of the system. In this paper, we propose a data-driven approach to load balance the I/O servers at scale, targeted at Lustre deployments. To this end, we design a global mapper on Lustre Metadata Server, which gathers runtime statistics from key storage components on the I/O path, and applies Markov chain modeling and a minimum-cost maximum-flow algorithm to decide where data should be placed. Evaluation using a realistic system simulator and a real setup shows that our approach yields better load balancing, which in turn can improve end-to-end performance.