Efficient constructions for large‐state block ciphers based on AES New Instructions

Efficient constructions for large‐state block ciphers based on AES New Instructions
复制标题

基于AES新指令的大状态分组密码的高效构建

DOI:
10.1049/ise2.12053
复制
发表时间:
2021
影响因子:
1.4
通讯作者:
Isobe Takanori
Isobe Takanori
中科院分区:
计算机科学4区
文献类型:
--
作者:
Shiba Rentaro;Sakamoto Kosei;Isobe Takanori

文献摘要

相似文献

从长期安全性的角度来看,具有256位或512位块大小的大状态分组密码受到了广泛关注。现有的大状态块密码,如Haraka-v2和Pholkos,仅由AES新指令集(AES-NI)和一个字混洗组成,可以通过SIMD指令有效地执行,以实现快速软件实现。在Haraka-v2和Pholkos中,AES轮函数在每个步骤中并行执行两次,其输出被混洗(称为两轮构造)。在这项研究中,基于AES-NI的最佳构造和高效的字混洗,这样的大状态块密码在软件加密速度方面进行了探索。具体地,识别出可以在更少轮数中实现安全性的最佳类别的字混洗,该最佳类别的字混洗可以在SIMD中有效地实现以有助于提高大状态块密码的性能。测量每个CPU架构的速度。因此,作者揭示了这样的构造,即在每个步骤中并行执行两轮AES轮函数,其输出被混洗(称为两轮构造),并且在所有Skylake架构或更高版本的CPU中都是最佳的。此外,作者揭示了在字混洗指令的速度方面存在明显的差异,即使它们理论上需要相同数量的周期。因此,作者阐明了考虑到这些差异,为每种架构的最佳建设。
Large‐state block ciphers with 256 bits or 512 bits block sizes receive much attention from the viewpoint of long‐term security. Existing large‐state block ciphers, such as Haraka‐v2 and Pholkos, consist of only the AES New Instructions set (AES‐NI) and a word shuffle that can be efficiently executed by SIMD instructions for fast software implementation. In Haraka‐v2 and Pholkos, the AES round function is executed twice in parallel at each step and its outputs are shuffled (called two‐round constructions). In this study, optimal constructions based on AES‐NI and efficient word shuffles for such large‐state block ciphers in terms of the encryption speed for software are explored. Specifically, an optimal class of word shuffles that can achieve security in a smaller number of rounds from the class of word shuffles that can be efficiently implemented in SIMD to contribute to the improvement of the performance of large‐state block ciphers is identified. Their speed for each CPU architecture is measured. As a result, the authors reveal the constructions such that two rounds of the AES round function is executed in parallel at each step and its outputs are shuffled (called two‐round constructions) and are optimal in all CPUs with Skylake architecture or later versions. Furthermore, the authors reveal that there is a clear difference in word shuffle instructions with respect to the speed, even if they theoretically require the same number of cycles. Consequently, the authors clarify the optimal construction for each architecture by taking these differences into consideration.