BachGAN: High-Resolution Image Synthesis From Salient Object Layout

BachGAN: High-Resolution Image Synthesis From Salient Object Layout
复制标题

DOI:
10.1109/cvpr42600.2020.00839
复制
发表时间:
2020-03
期刊:
2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Yandong Li;Yu Cheng;Zhe Gan;Licheng Yu;Liqiang Wang;Jingjing Liu
Yandong Li;Yu Cheng;Zhe Gan;Licheng Yu;Liqiang Wang;Jingjing Liu
中科院分区:
其他
文献类型:
--
作者:
Yandong Li;Yu Cheng;Zhe Gan;Licheng Yu;Liqiang Wang;Jingjing Liu

文献摘要

被引文献

相似文献

我们针对图像生成的更实际应用提出了一项新任务——从显著物体布局进行高质量图像合成。这种新设定要求用户仅提供显著物体的布局(即前景边界框和类别),并让模型用虚构的背景和匹配的前景完成绘制。这项新任务带来了两个主要挑战:(i)在没有分割图输入的情况下如何生成细粒度的细节和逼真的纹理;(ii)如何以无缝的方式创建背景并将其与独立物体融合。为了解决这些问题,我们提出了背景幻觉生成对抗网络(BachGAN),它利用一个背景检索模块首先从一个大型候选池中选择一组分割图,然后通过一个背景融合模块对这些候选布局进行编码,为给定物体虚构一个合适的背景。通过动态生成虚构的背景表示,我们的模型能够合成具有逼真前景和完整背景的高分辨率图像。在Cityscapes和ADE20K数据集上的实验证明了BachGAN相对于现有方法的优势,这体现在生成图像的视觉保真度以及输出图像和输入布局之间的视觉对齐上。
We propose a new task towards more practical applications for image generation - high-quality image synthesis from salient object layout. This new setting requires users to provide only the layout of salient objects (i.e., foreground bounding boxes and categories) and lets the model complete the drawing with an invented background and a matching foreground. Two main challenges spring from this new task: (i) how to generate fine-grained details and realistic textures without segmentation map input; and (ii) how to create and weave a background into standalone objects in a seamless way. To tackle this, we propose Background Hallucination Generative Adversarial Network (BachGAN), which leverages a background retrieval module to first select a set of segmentation maps from a large candidate pool, then encodes these candidate layouts via a background fusion module to hallucinate a suitable background for the given objects. By generating the hallucinated background representation dynamically, our model can synthesize high-resolution images with both photo-realistic foreground and integral background. Experiments on Cityscapes and ADE20K datasets demonstrate the advantage of BachGAN over existing approaches, measured on both visual fidelity of generated images and visual alignment between output images and input layouts.