Remember the Past: Distilling Datasets into Addressable Memories for Neural Networks

Remember the Past: Distilling Datasets into Addressable Memories for Neural Networks
复制标题

DOI:
10.48550/arxiv.2206.02916
复制
发表时间:
2022-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Zhiwei Deng-;Olga Russakovsky
Zhiwei Deng-;Olga Russakovsky
中科院分区:
其他
文献类型:
--
作者:
Zhiwei Deng-;Olga Russakovsky

文献摘要

被引文献

相似文献

我们提出一种算法,该算法将大型数据集的关键信息压缩到紧凑的可寻址存储器中。然后可以调用这些存储器来快速重新训练神经网络并恢复性能(而不是在完整的原始数据集上进行存储和重新训练)。在数据集蒸馏框架的基础上,我们有一个关键的观察结果,即共享的通用表示允许更高效和有效的蒸馏。具体来说,我们学习一组基(也称为“存储器”),这些基在类别之间共享,并通过学习到的灵活寻址函数组合,以生成多样化的训练示例集。这带来了几个好处:1)压缩数据的大小不一定随类别数量线性增长;2)实现了更高的总体压缩率以及更有效的蒸馏;3)除了召回原始类别之外,还允许更通用的查询。我们在六个基准数据集上展示了数据集蒸馏任务的最先进结果,包括在蒸馏CIFAR10和CIFAR100时,保留精度分别提高了高达16.5%和9.7%。然后,我们利用我们的框架进行持续学习,在四个基准上取得了最先进的结果,在MANY上精度提高了23.2%。代码发布在我们的项目网页https://github.com/princetonvisualai/RememberThePast - DatasetDistillation上。
We propose an algorithm that compresses the critical information of a large dataset into compact addressable memories. These memories can then be recalled to quickly re-train a neural network and recover the performance (instead of storing and re-training on the full original dataset). Building upon the dataset distillation framework, we make a key observation that a shared common representation allows for more efficient and effective distillation. Concretely, we learn a set of bases (aka ``memories'') which are shared between classes and combined through learned flexible addressing functions to generate a diverse set of training examples. This leads to several benefits: 1) the size of compressed data does not necessarily grow linearly with the number of classes; 2) an overall higher compression rate with more effective distillation is achieved; and 3) more generalized queries are allowed beyond recalling the original classes. We demonstrate state-of-the-art results on the dataset distillation task across six benchmarks, including up to 16.5% and 9.7% in retained accuracy improvement when distilling CIFAR10 and CIFAR100 respectively. We then leverage our framework to perform continual learning, achieving state-of-the-art results on four benchmarks, with 23.2% accuracy improvement on MANY. The code is released on our project webpage https://github.com/princetonvisualai/RememberThePast-DatasetDistillation.