Controlling StyleGANs using rough scribbles via one‐shot learning

Controlling StyleGANs using rough scribbles via one‐shot learning
复制标题

DOI:
10.1002/cav.2102
复制
发表时间:
2022-07
影响因子:
1.1
通讯作者:
Yuki Endo;Yoshihiro Kanamori
Yuki Endo;Yoshihiro Kanamori
中科院分区:
计算机科学4区
文献类型:
--
作者:
Yuki Endo;Yoshihiro Kanamori

文献摘要

相似文献

本文解决了从粗糙的稀疏注释中进行一次性语义图像合成的挑战性问题,我们称之为“语义涂鸦”。也就是说,仅从用语义涂鸦注释的单个训练对中,我们就可以生成逼真且多样化的图像,并可对面部部位布局和身体姿势等进行布局控制。 We present a training strategy that performs pseudo labeling for semantic scribbles using the StyleGAN prior. Our key idea is to construct a simple mapping between StyleGAN features and each semantic class from a single example of semantic scribbles.通过这样的映射,我们可以从随机噪声中生成无限数量的伪语义涂鸦,以训练编码器来控制预训练的 StyleGAN 生成器。即使我们通过一次性监督获得了粗糙的伪语义涂鸦,由于我们的 GAN 反转框架,我们的方法也可以合成高质量的图像。 We further offer optimization‐based postprocessing to refine the pixel alignment of synthesized images. Qualitative and quantitative results on various datasets demonstrate improvement over previous approaches in one‐shot settings.
This paper tackles the challenging problem of one‐shot semantic image synthesis from rough sparse annotations, which we call “semantic scribbles.” Namely, from only a single training pair annotated with semantic scribbles, we generate realistic and diverse images with layout control over, for example, facial part layouts and body poses. We present a training strategy that performs pseudo labeling for semantic scribbles using the StyleGAN prior. Our key idea is to construct a simple mapping between StyleGAN features and each semantic class from a single example of semantic scribbles. With such mappings, we can generate an unlimited number of pseudo semantic scribbles from random noise to train an encoder for controlling a pretrained StyleGAN generator. Even with our rough pseudo semantic scribbles obtained via one‐shot supervision, our method can synthesize high‐quality images thanks to our GAN inversion framework. We further offer optimization‐based postprocessing to refine the pixel alignment of synthesized images. Qualitative and quantitative results on various datasets demonstrate improvement over previous approaches in one‐shot settings.