lunule: An Agile and Judicious Metadata Load Balancer for CephFS

lunule: An Agile and Judicious Metadata Load Balancer for CephFS
复制标题

DOI:
10.1145/3458817.3476196
复制
发表时间:
2021-11
期刊:
SC21: International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子:
--
通讯作者:
Yiduo Wang;Cheng Li;Xinyang Shao;Youxu Chen;Feng Yan;Yinlong Xu
Yiduo Wang;Cheng Li;Xinyang Shao;Youxu Chen;Feng Yan;Yinlong Xu
中科院分区:
其他
文献类型:
--
作者:
Yiduo Wang;Cheng Li;Xinyang Shao;Youxu Chen;Feng Yan;Yinlong Xu

文献摘要

相似文献

十年来,Ceph分布式文件系统(CephFS)已被广泛应用于服务从互联网服务到人工智能计算等许多关键领域不断增长的大数据。为了横向扩展海量元数据访问,CephFS采用动态子树分区方法,分割分层命名空间并将子树分布在多个元数据服务器上。然而,该方法存在严重的不平衡问题,由于不准确的不平衡预测、对工作负载特征的忽视以及不必要/无效的迁移活动,可能会导致性能不佳。为了消除这些低效率,我们提出了 Lunule,一种新颖的 CephFS 元数据负载均衡器,它采用不平衡因子模型来准确确定何时触发重新平衡并容忍良性不平衡情况。 Lunule 进一步采用工作负载感知迁移规划器来适当选择子树迁移候选者。与基线相比,Lunule 实现了更好的负载平衡,对于五个实际工作负载及其混合,分别将元数据吞吐量提高了 315.8%,并将尾部作业完成时间缩短了 64.6%。此外,Lunule 能够处理元数据集群扩展和客户端工作负载增长,并在 16 个 MDS 的集群上线性扩展。
For a decade, the Ceph distributed file system (CephFS) has been widely used to serve the ever-growing big data in many key fields ranging from Internet services to AI computing. To scale out the massive metadata access, CephFS adopts a dynamic subtree partitioning method, splitting the hierarchical namespace and distributing subtrees across multiple metadata servers. However, this method suffers from a severe imbalance problem that may result in poor performance due to its inaccurate imbalance prediction, ignorance of workload characteristics, and unnecessary/invalid migration ac-tivities. To eliminate these inefficiencies, we propose Lunule, a novel CephFS metadata load balancer, which employs an imbalance fac-tor model for accurately determining when to trigger re-balance and tolerate benign imbalanced situations. Lunule further adopts a workload-aware migration planner to appropriately select sub-tree migration candidates. Compared to baselines, Lunule achieves better load balance, increases the metadata throughput by up to 315.8%, and shortens the tail job completion time by up to 64.6% for five real-world workloads and their mixture, respectively. Be-sides, Lunule is capable of handling the metadata cluster expansion and the client workload growth, and scales linearly on a cluster of 16 MDSs.