Scalable and adaptive log manager in distributed systems

Scalable and adaptive log manager in distributed systems
复制标题

DOI:
10.1007/s11704-022-1357-5
复制
发表时间:
2022-08
影响因子:
4.2
通讯作者:
Huan Zhou;Weining Qian;Xuan Zhou;Qiwen Dong;Aoying Zhou;Wenrong Tan
Huan Zhou;Weining Qian;Xuan Zhou;Qiwen Dong;Aoying Zhou;Wenrong Tan
中科院分区:
计算机科学3区
文献类型:
--
作者:
Huan Zhou;Weining Qian;Xuan Zhou;Qiwen Dong;Aoying Zhou;Wenrong Tan

文献摘要

相似文献

联机事务处理(OLTP)系统依靠事务日志和基于仲裁的共识协议来保证持久性、高可用性和强一致性。这使得日志管理器成为分布式数据库管理系统(ddbms)的关键组件。ddbms的leader通常采用集中式日志方式将日志条目写入稳定的存储设备,并使用恒定的日志复制策略定期将其状态同步到follower。随着新硬件的出现和事务处理的高并行性,传统的集中式日志设计限制了可扩展性,并且恒定的复制触发条件不能在动态工作负载下始终保持最佳性能。在本文中,我们提出了一个新的日志管理器,名为salmoo,具有可扩展的日志记录和自适应复制功能,适用于分布式数据库系统。可扩展日志通过利用高度并发的数据结构和快速的日志漏洞跟踪消除了集中式争用。自适应复制的核心是一种自适应日志传输方法,该方法根据实时工作负载动态调整在leader和follower之间传输的日志条目数。我们在开源事务处理系统Cedar和DBx1000中实现并评估了Salmo。实验结果表明,Salmo通过增加工作线程数实现了良好的扩展性,与Raft的日志复制相比,峰值吞吐量提高了1.56倍,延迟降低了4倍以上,并且在动态工作负载下始终保持高效稳定的性能。
On-line transaction processing (OLTP) systems rely on transaction logging and quorum-based consensus protocol to guarantee durability, high availability and strong consistency. This makes the log manager a key component of distributed database management systems (DDBMSs). The leader of DDBMSs commonly adopts a centralized logging method to writing log entries into a stable storage device and uses a constant log replication strategy to periodically synchronize its state to followers. With the advent of new hardware and high parallelism of transaction processing, the traditional centralized design of logging limits scalability, and the constant trigger condition of replication can not always maintain optimal performance under dynamic workloads.In this paper, we propose a new log manager namedSalmowith scalable logging and adaptive replication for distributed database systems. The scalable logging eliminates centralized contention by utilizing a highly concurrent data structure and speedy log hole tracking. The kernel of adaptive replication is an adaptive log shipping method, which dynamically adjusts the number of log entries transmitted between leader and followers based on the real-time workload. We implemented and evaluated Salmo in the open-sourced transaction processing systems Cedar and DBx1000. Experimental results show that Salmo scales well by increasing the number of working threads, improves peak throughput by 1.56× and reduces latency by more than 4× over log replication of Raft, and maintains efficient and stable performance under dynamic workloads all the time.