lunule: An Agile and Judicious Metadata Load Balancer for CephFS
lunule: An Agile and Judicious Metadata Load Balancer for CephFS
复制标题
DOI:
10.1145/3458817.3476196
复制
发表时间:
2021-11
期刊:
影响因子:
--
通讯作者:
Yiduo Wang;Cheng Li;Xinyang Shao;Youxu Chen;Feng Yan;Yinlong Xu
中科院分区:
文献类型:
--
作者:
Yiduo Wang;Cheng Li;Xinyang Shao;Youxu Chen;Feng Yan;Yinlong Xu
For a decade, the Ceph distributed file system (CephFS) has been widely used to serve the ever-growing big data in many key fields ranging from Internet services to AI computing. To scale out the massive metadata access, CephFS adopts a dynamic subtree partitioning method, splitting the hierarchical namespace and distributing subtrees across multiple metadata servers. However, this method suffers from a severe imbalance problem that may result in poor performance due to its inaccurate imbalance prediction, ignorance of workload characteristics, and unnecessary/invalid migration ac-tivities. To eliminate these inefficiencies, we propose Lunule, a novel CephFS metadata load balancer, which employs an imbalance fac-tor model for accurately determining when to trigger re-balance and tolerate benign imbalanced situations. Lunule further adopts a workload-aware migration planner to appropriately select sub-tree migration candidates. Compared to baselines, Lunule achieves better load balance, increases the metadata throughput by up to 315.8%, and shortens the tail job completion time by up to 64.6% for five real-world workloads and their mixture, respectively. Be-sides, Lunule is capable of handling the metadata cluster expansion and the client workload growth, and scales linearly on a cluster of 16 MDSs.