Memory Efficient Diffusion Probabilistic Models via Patch-based Generation

Memory Efficient Diffusion Probabilistic Models via Patch-based Generation
复制标题

DOI:
10.48550/arxiv.2304.07087
复制
发表时间:
2023-04
期刊:
ArXiv
影响因子:
--
通讯作者:
Shinei Arakawa;Hideki Tsunashima;Daichi Horita;Keitaro Tanaka;S. Morishima
Shinei Arakawa;Hideki Tsunashima;Daichi Horita;Keitaro Tanaka;S. Morishima
中科院分区:
其他
文献类型:
--
作者:
Shinei Arakawa;Hideki Tsunashima;Daichi Horita;Keitaro Tanaka;S. Morishima

文献摘要

相似文献

扩散概率模型已成功生成高质量且多样化的图像。然而,传统模型的输入和输出都是高分辨率图像,对内存的要求过高,这使得它们对边缘设备不太实用。以前的生成对抗网络方法提出了一种基于补丁的方法,该方法使用位置编码和全局内容信息。然而,设计一个基于补丁的方法扩散概率模型是不平凡的。在本文中,我们重新发送一个扩散概率模型,生成图像上的补丁补丁。我们提出了两个条件的方法补丁为基础的一代。首先,我们提出了位置方面的条件反射使用一个热表示,以确保补丁是在正确的位置。其次,我们提出了全局内容调节(GCC),以确保补丁连接在一起时具有连贯的内容。我们在CelebA和LSUN卧室数据集上定性和定量地评估了我们的模型,并证明了最大内存消耗和生成的图像质量之间的适度权衡。具体来说,当整个图像被分成2 × 2块时,我们提出的方法可以将最大内存消耗减少一半,同时保持相当的图像质量。
Diffusion probabilistic models have been successful in generating high-quality and diverse images. However, traditional models, whose input and output are high-resolution images, suffer from excessive memory requirements, making them less practical for edge devices. Previous approaches for generative adversarial networks proposed a patch-based method that uses positional encoding and global content information. Nevertheless, designing a patch-based approach for diffusion probabilistic models is non-trivial. In this paper, we resent a diffusion probabilistic model that generates images on a patch-by-patch basis. We propose two conditioning methods for a patch-based generation. First, we propose position-wise conditioning using one-hot representation to ensure patches are in proper positions. Second, we propose Global Content Conditioning (GCC) to ensure patches have coherent content when concatenated together. We evaluate our model qualitatively and quantitatively on CelebA and LSUN bedroom datasets and demonstrate a moderate trade-off between maximum memory consumption and generated image quality. Specifically, when an entire image is divided into 2 x 2 patches, our proposed approach can reduce the maximum memory consumption by half while maintaining comparable image quality.