Efficient Cache Utilization via Model-aware Data Placement for Recommendation Models

Efficient Cache Utilization via Model-aware Data Placement for Recommendation Models
复制标题

通过推荐模型的模型感知数据放置实现高效缓存利用

DOI:
--
复制
发表时间:
2021
期刊:
International Symposium on Memory Systems
影响因子:
--
通讯作者:
Shaizeen Aga
Shaizeen Aga
中科院分区:
--
文献类型:
--
作者:
M. Ibrahim;Onur Kayiran;Shaizeen Aga

文献摘要

被引文献

相似文献

基于深度神经网络(DNN)的推荐模型(RM)代表了一类广泛用于社交媒体、娱乐内容和在线业务的关键工作负载。考虑到它们的普遍使用,理解这些模型的内存子系统行为是至关重要的,特别是从未来内存子系统设计的角度来看。为此,在这项工作中,我们首先做了深入的内存足迹和流量分析新兴的RM。我们观察到,新兴的RM将严重强调未来(可能更大)的缓存和内存。为了解决这一挑战,我们提出了一个关键的观察,即一个数据放置策略,是知道这些模型中的组件(而不是一个认为整个模型作为一个整体)站在一个更好的机会,减轻压力的内存子系统。具体来说,这些模型的两个关键组成部分,即嵌入表和多层感知器层,我们展示了如何利用内存访问嵌入表的局部性来提出一个更细致入微的数据放置方案。我们展示了我们提出的数据放置策略如何减少整体内存流量(约32%),同时提高性能(高达1.99倍)。我们认为,内存子系统,更适合驻留控制站在一个更好的机会,以满足新兴模式的需求。
Deep neural network (DNN) based recommendation models (RMs) represent a class of critical workloads that are broadly used in social media, entertainment content, and online businesses. Given their pervasive usage, understanding the memory subsystem behavior of these models is crucial, particularly from the perspective of future memory subsystem design. To this end, in this work, we first do an in-depth memory footprint and traffic analysis of emerging RMs. We observe that emerging RMs will severely stress future (and possibly larger) caches and memories. To address this challenge, we make the key observation that a data placement strategy that is aware of the components within these models (as opposed to one that considers the entire model as a whole) stands a better chance of relieving the stress on the memory subsystem. Specifically, of the two key components of these models, namely, embedding tables and multi-layer perceptron layers, we show how we can exploit the locality of memory accesses to embedding tables to come up with a more nuanced data placement scheme. We demonstrate how our proposed data placement strategy can reduce overall memory traffic (approximately 32%) while improving performance (up to 1.99 ×). We argue that memory subsystems that are more amenable to residency controls stand a better chance to address the needs of emerging models.