Efficient Cache Utilization via Model-aware Data Placement for Recommendation Models
Efficient Cache Utilization via Model-aware Data Placement for Recommendation Models
复制标题
通过推荐模型的模型感知数据放置实现高效缓存利用
DOI:
--
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
Shaizeen Aga
中科院分区:
文献类型:
--
作者:
M. Ibrahim;Onur Kayiran;Shaizeen Aga
Deep neural network (DNN) based recommendation models (RMs) represent a class of critical workloads that are broadly used in social media, entertainment content, and online businesses. Given their pervasive usage, understanding the memory subsystem behavior of these models is crucial, particularly from the perspective of future memory subsystem design. To this end, in this work, we first do an in-depth memory footprint and traffic analysis of emerging RMs. We observe that emerging RMs will severely stress future (and possibly larger) caches and memories. To address this challenge, we make the key observation that a data placement strategy that is aware of the components within these models (as opposed to one that considers the entire model as a whole) stands a better chance of relieving the stress on the memory subsystem. Specifically, of the two key components of these models, namely, embedding tables and multi-layer perceptron layers, we show how we can exploit the locality of memory accesses to embedding tables to come up with a more nuanced data placement scheme. We demonstrate how our proposed data placement strategy can reduce overall memory traffic (approximately 32%) while improving performance (up to 1.99 ×). We argue that memory subsystems that are more amenable to residency controls stand a better chance to address the needs of emerging models.