iez: Resource Contention Aware Load Balancing for Large-Scale Parallel File Systems

iez: Resource Contention Aware Load Balancing for Large-Scale Parallel File Systems
复制标题

DOI:
10.1109/ipdps.2019.00070
复制
发表时间:
2019-05
期刊:
2019 IEEE International Parallel and Distributed Processing Symposium (IPDPS)
影响因子:
--
通讯作者:
Bharti Wadhwa;A. Paul;Sarah Neuwirth;Feiyi Wang;S. Oral;A. Butt;Jon Bernard;K. Cameron
Bharti Wadhwa;A. Paul;Sarah Neuwirth;Feiyi Wang;S. Oral;A. Butt;Jon Bernard;K. Cameron
中科院分区:
其他
文献类型:
--
作者:
Bharti Wadhwa;A. Paul;Sarah Neuwirth;Feiyi Wang;S. Oral;A. Butt;Jon Bernard;K. Cameron

文献摘要

被引文献

相似文献

并行I/O性能对于在大规模高性能计算(HPC)系统上支持科学应用至关重要。但是,底层分布式和共享存储系统中的I/O负载不平衡会显著降低整体应用程序性能。缓解这种负载失衡有两个相互冲突的挑战:(I)优化系统范围的数据布局以最大化分布式存储服务器的带宽优势,即跨应用程序和作业运行高效地分配I/O资源;(Ii)优化以客户端为中心的数据移动以最小化客户端和服务器之间的I/O负载请求延迟,即高效地将服务中的I/O资源分配给单个应用程序和作业运行。此外,需要更改应用程序的现有方法限制了在商业或专有部署中的广泛采用。我们提出了IEZ,这是一个“端到端的控制平面”,其中客户端透明且自适应地写入一组选定的I/O服务器,以实现平衡的数据放置。我们的控制平面利用实时负载信息进行分布式存储服务器全局数据放置,而我们的设计模型利用基于跟踪的优化技术来最小化客户端和服务器之间的I/O负载请求延迟。我们在一个实验集群上对两个常见的用例进行了评估:大型顺序写入的合成I/O基准IOR和科学应用程序I/O内核HACC-I/O。结果显示,与现有技术相比,读写性能分别提高了34%和32%。
Parallel I/O performance is crucial to sustaining scientific applications on large-scale High-Performance Computing (HPC) systems. However, I/O load imbalance in the underlying distributed and shared storage systems can significantly reduce overall application performance. There are two conflicting challenges to mitigate this load imbalance: (i) optimizing systemwide data placement to maximize the bandwidth advantages of distributed storage servers, i.e., allocating I/O resources efficiently across applications and job runs; and (ii) optimizing client-centric data movement to minimize I/O load request latency between clients and servers, i.e., allocating I/O resources efficiently in service to a single application and job run. Moreover, existing approaches that require application changes limit wide-spread adoption in commercial or proprietary deployments. We propose iez, an "end-to-end control plane" where clients transparently and adaptively write to a set of selected I/O servers to achieve balanced data placement. Our control plane leverages realtime load information for distributed storage server global data placement while our design model leverages trace-based optimization techniques to minimize I/O load request latency between clients and servers. We evaluate our proposed system on an experimental cluster for two common use cases: synthetic I/O benchmark IOR for large sequential writes and a scientific application I/O kernel, HACC-I/O. Results show read and write performance improvements of up to 34% and 32%, respectively, compared to the state of the art.