Characterizing and optimizing TPC-C workloads on large-scale systems using SSD arrays

Characterizing and optimizing TPC-C workloads on large-scale systems using SSD arrays
复制标题

使用 SSD 阵列表征和优化大型系统上的 TPC-C 工作负载

DOI:
10.1007/s11432-015-5383-x
复制
发表时间:
2016
期刊:
Science China Information Sciences
影响因子:
--
通讯作者:
Zheng Weimin
Zheng Weimin
中科院分区:
其他
文献类型:
--
作者:
Zhai Jidong;Zhang Feng;Li Qingwen;Chen Wenguang;Zheng Weimin

文献摘要

相似文献

事务处理性能基准测试理事会C(TPC-C)是评估运行联机事务处理应用程序的高端计算机性能的事实上的标准。与其他标准基准不同,事务处理性能理事会只定义了TPC-C基准的规范,但没有为最终用户提供任何标准实现。由于TPC-C工作负载的复杂性,在大规模高端计算机上获得TPC-C评估的最佳性能是一项具有挑战性的任务。在本文中,我们设计并实现了一个大规模的TPC-C评估系统的基础上,最新的TPC-C规范使用固态硬盘(SSD)存储设备。通过分析TPC-C工作负载的特点,提出了一系列系统级优化方法来提高TPC-C的性能。首先,我们提出了一种基于SmallFile表空间的方法来组织测试数据在所有的磁盘阵列分区的循环方法,这可以充分利用底层的磁盘阵列。其次,我们提出了一个基于NOOP的磁盘调度算法,以降低处理器的利用率,提高平均输入/输出服务时间。第三,为了提高系统翻译后备缓冲区的命中率,减少处理器开销,我们利用巨大的页面技术来管理大量的内存资源。最后,我们根据非均匀内存访问系统的非对称性特征,提出了一种位置感知的中断映射策略,以提高系统性能。使用这些优化方法,我们在两台使用SSD阵列的大型高端计算机上进行了TPC-C测试。实验结果表明,该方法能有效地提高TPC-C的性能。例如,在英特尔韦斯特米尔服务器上进行的TPC-C测试的性能达到了每分钟101.8万次交易。
Transaction processing performance council benchmark C (TPC-C) is the de facto standard for evaluating the performance of high-end computers running on-line transaction processing applications. Differing from other standard benchmarks, the transaction processing performance council only defines specifications for the TPC-C benchmark, but does not provide any standard implementation for end-users. Due to the complexity of the TPC-C workload, it is a challenging task to obtain optimal performance for TPC-C evaluation on a large-scale high-end computer. In this paper, we designed and implemented a large-scale TPC-C evaluation system based on the latest TPC-C specification using solid-state drive (SSD) storage devices. By analyzing the characteristics of the TPC-C workload, we propose a series of system-level optimization methods to improve the TPC-C performance. First, we propose an approach based on SmallFile table space to organize the test data in a round-robin method on all of the disk array partitions; this can make full use of the underlying disk arrays. Second, we propose using a NOOP-based disk scheduling algorithm to reduce the utilization rate of processors and improve the average input/output service time. Third, to improve the system translation lookaside buffer hit rate and reduce the processor overhead, we take advantage of the huge page technique to manage a large amount of memory resources. Lastly, we propose a locality-aware interrupt mapping strategy based on the asymmetry characteristic of non-uniform memory access systems to improve the system performance. Using these optimization methods, we performed the TPC-C test on two large-scale high-end computers using SSD arrays. The experimental results show that our methods can effectively improve the TPC-C performance. For example, the performance of the TPC-C test on an Intel Westmere server reached 1.018 million transactions per minute.