Learning Layout and Style Reconfigurable GANs for Controllable Image Synthesis

Learning Layout and Style Reconfigurable GANs for Controllable Image Synthesis
复制标题

用于可控图像合成的学习布局和风格可重构GANN

DOI:
10.1109/tpami.2021.3078577
复制
发表时间:
2022-09-01
影响因子:
23.6
通讯作者:
Wu, Tianfu
Wu, Tianfu
中科院分区:
计算机科学1区
文献类型:
--
作者:
Sun, Wei;Wu, Tianfu

文献摘要

被引文献

相似文献

随着最近在学习深度生成模型方面取得的显著进展,从可重构结构化输入开发可控图像合成模型变得越来越有趣。本文重点关注最近出现的任务,布局到图像,其目标是学习生成模型,用于从空间布局合成照片级逼真的图像(即,配置在图像网格中的对象边界框)及其样式代码(即,由潜在向量编码的结构和外观变化)。本文首先提出了一个直观的范式的任务,布局到掩模到图像,学习展开对象掩模在弱监督的方式基于输入布局和对象样式代码。布局到掩模组件与生成器网络中的层深度交互,以桥接输入布局和合成图像之间的差距。然后,本文提出了一种基于生成对抗网络(GAN)的方法,用于拟议的布局到掩模到图像合成,并在图像和对象级别进行布局和风格控制。可控性实现了一个新的实例敏感和布局感知归一化(ISLA-Norm)计划。在不牺牲性能的情况下,进一步开发了所提出方法的布局半监督版本。在实验中,所提出的方法在COCO-Stuff数据集和Visual Genome数据集上进行了测试,获得了最先进的性能。
With the remarkable recent progress on learning deep generative models, it becomes increasingly interesting to develop models for controllable image synthesis from reconfigurable structured inputs. This paper focuses on a recently emerged task, layout-to-image, whose goal is to learn generative models for synthesizing photo-realistic images from a spatial layout (i.e., object bounding boxes configured in an image lattice) and its style codes (i.e., structural and appearance variations encoded by latent vectors). This paper first proposes an intuitive paradigm for the task, layout-to-mask-to-image, which learns to unfold object masks in a weakly-supervised way based on an input layout and object style codes. The layout-to-mask component deeply interacts with layers in the generator network to bridge the gap between an input layout and synthesized images. Then, this paper presents a method built on Generative Adversarial Networks (GANs) for the proposed layout-to-mask-to-image synthesis with layout and style control at both image and object levels. The controllability is realized by a proposed novel Instance-Sensitive and Layout-Aware Normalization (ISLA-Norm) scheme. A layout semi-supervised version of the proposed method is further developed without sacrificing performance. In experiments, the proposed method is tested in the COCO-Stuff dataset and the Visual Genome dataset with state-of-the-art performance obtained.