LAORAM: A Look Ahead ORAM Architecture for Training Large Embedding Tables

LAORAM: A Look Ahead ORAM Architecture for Training Large Embedding Tables
复制标题

DOI:
10.1145/3579371.3589111
复制
发表时间:
2021-07
期刊:
Proceedings of the 50th Annual International Symposium on Computer Architecture
影响因子:
--
通讯作者:
Rachit Rajat;Yongqin Wang;M. Annavaram
Rachit Rajat;Yongqin Wang;M. Annavaram
中科院分区:
其他
文献类型:
--
作者:
Rachit Rajat;Yongqin Wang;M. Annavaram

文献摘要

相似文献

内存访问模式已经被证明会泄漏关键信息,如安全密钥和程序的空间和时间信息。这种信息泄漏对具有嵌入表的机器学习模型构成了重大的隐私挑战。嵌入表用于从训练数据中学习分类特征。嵌入表条目的地址携带隐私敏感信息,因为条目的地址公开了与用户相关联的特征。不经意RAM(ORAM)及其增强型变体(如PathORAM)已成为隐藏内存访问流泄漏的可行解决方案。PathORAM为每个内存提取请求提取整个内存块路径,从而导致大量带宽和性能开销。在这项工作中,我们提出了前瞻ORAM(LAORAM),一个ORAM框架,旨在保护用户隐私嵌入表训练。LAORAM利用了ML训练的独特属性,即未来将要使用的训练样本是事先已知的。LAORAM对训练样本进行预处理,以识别在不久的将来一起访问的内存块。LAORAM将一起访问的多个块组合为超级块,并尝试将超级块中的所有块分配给少数路径。因此,未来对块集合的访问可以从几条路径满足,有效地减少了对ORAM的读取和写入的数量。为了进一步提高性能,LAORAM为PathORAM使用了胖树结构,即具有可变桶大小的树,有效地减少了所需的后台驱逐次数,从而提高了存储使用率。我们已经评估了LAORAM使用推荐模型(DLRM)和NLP模型(XLM-R)嵌入表配置。LAORAM在推荐数据集(Kaggle)上的执行速度比PathORAM快5倍,在NLP数据集(XNLI)上快5.4倍,同时保证与原始PathORAM相同的安全保证。
Memory access patterns have been demonstrated to leak critical information such as security keys and a program's spatial and temporal information. This information leak poses a significant privacy challenge in machine learning models with embedding tables. Embedding tables are used to learn categorical features from training data. The address of an embedding table entry carries privacy sensitive information since the address of an entry discloses features associated with a user. Oblivious RAM (ORAM), and its enhanced variants, such as PathORAM, have emerged as viable solutions to hide leakage from memory access streams. PathORAM fetches an entire path of memory blocks for every memory fetch request, thereby leading to substantial bandwidth and performance overheads. In this work, we present Look Ahead ORAM (LAORAM), an ORAM framework designed to protect user privacy during embedding table training. LAORAM exploits the unique property of ML training, namely the training samples that are going to be used in the future are known beforehand. LAORAM preprocesses the training samples to identify the memory blocks which are accessed together in the near future. LAORAM combines multiple blocks accessed together as superblocks and tries to assign all blocks in a superblock to few paths. Thus, future accesses to a collection of blocks can be satisfied from a few paths, effectively reducing the number of reads and writes to the ORAM. To further increase performance, LAORAM uses a fat-tree structure for PathORAM, i.e. a tree with variable bucket size, effectively reducing the number of background evictions required, which improves the stash usage. We have evaluated LAORAM using both a recommendation model (DLRM) and an NLP model (XLM-R) embedding table configurations. LAORAM performs 5 times faster than PathORAM on a recommendation dataset (Kaggle) and 5.4 times faster on an NLP dataset (XNLI) while guaranteeing the same security guarantees as the original PathORAM.