ProSE: the architecture and design of a protein discovery engine

ProSE: the architecture and design of a protein discovery engine
复制标题

DOI:
10.1145/3503222.3507722
复制
发表时间:
2022-02
期刊:
Proceedings of the 27th ACM International Conference on Architectural Support for Programming Languages and Operating Systems
影响因子:
--
通讯作者:
Eyes S. Robson;Ceyu Xu;Lisa Wu Wills
Eyes S. Robson;Ceyu Xu;Lisa Wu Wills
中科院分区:
其他
文献类型:
--
作者:
Eyes S. Robson;Ceyu Xu;Lisa Wu Wills

文献摘要

相似文献

蛋白质语言模型为蛋白质结构预测、功能注释和药物发现提供了突破性的方法。广泛采用这些强大模型的主要限制是与这些模型的训练和推理相关的高计算成本,特别是在较长的序列长度下。我们提出的架构,微架构,和硬件实现的蛋白质设计和发现加速器,ProSE(蛋白质syslogengine)。ProSE拥有一系列定制的异构脉动阵列和特殊功能,可以有效地处理迁移学习模型的推理。该架构将SIMD风格的计算与脉动阵列架构相结合,优化了跨模型层的粗粒度操作序列,以在不牺牲通用性的情况下实现效率。ProSE执行蛋白质BERT推理的加速比高达6.9倍,能效(性能/瓦特)为NVIDIA A100 GPU的48倍。与TPUv 3(TPUv 2)相比,ProSE实现了高达5.5倍(12.7倍)的加速比和173倍(249倍)的能效。
Protein language models have enabled breakthrough approaches to protein structure prediction, function annotation, and drug discovery. A primary limitation to the widespread adoption of these powerful models is the high computational cost associated with the training and inference of these models, especially at longer sequence lengths. We present the architecture, microarchitecture, and hardware implementation of a protein design and discovery accelerator, ProSE (Protein Systolic Engine). ProSE has a collection of custom heterogeneous systolic arrays and special functions that process transfer learning model inferences efficiently. The architecture marries SIMD-style computations with systolic array architectures, optimizing coarse-grained operation sequences across model layers to achieve efficiency without sacrificing generality. ProSE performs Protein BERT inference at up to 6.9× speedup and 48× power efficiency (performance/Watt) compared to one NVIDIA A100 GPU. ProSE achieves up to 5.5 × (12.7×) speedup and 173× (249×) power efficiency compared to TPUv3 (TPUv2).