MixNMatch: Multifactor Disentanglement and Encoding for Conditional Image Generation

MixNMatch: Multifactor Disentanglement and Encoding for Conditional Image Generation
复制标题

DOI:
10.1109/cvpr42600.2020.00806
复制
发表时间:
2019-11
期刊:
2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Yuheng Li;Krishna Kumar Singh;Utkarsh Ojha;Yong Jae Lee
Yuheng Li;Krishna Kumar Singh;Utkarsh Ojha;Yong Jae Lee
中科院分区:
其他
文献类型:
--
作者:
Yuheng Li;Krishna Kumar Singh;Utkarsh Ojha;Yong Jae Lee

文献摘要

相似文献

我们提出了MixNMatch,一个有条件的生成模型,学习解开和编码的背景,对象的姿势,形状和纹理从真实的图像与最小的监督,混合和匹配图像生成。我们建立在FineGAN(一种无条件生成模型)的基础上,学习所需的解纠缠和图像生成器,并利用对抗性联合图像代码分布匹配来学习潜在因子编码器。MixNMatch在训练过程中需要边界框来建模背景,但不需要其他监督。通过大量的实验,我们证明了MixNMatch的能力,准确地解开,编码,并联合收割机多个因素的混合和匹配的图像生成,包括sketch 2color,cartoon 2 img和img 2gif应用程序。我们的代码/模型/演示可以在https://github.com/Yuheng-Li/MixNMatch上找到
We present MixNMatch, a conditional generative model that learns to disentangle and encode background, object pose, shape, and texture from real images with minimal supervision, for mix-and-match image generation. We build upon FineGAN, an unconditional generative model, to learn the desired disentanglement and image generator, and leverage adversarial joint image-code distribution matching to learn the latent factor encoders. MixNMatch requires bounding boxes during training to model background, but requires no other supervision. Through extensive experiments, we demonstrate MixNMatch's ability to accurately disentangle, encode, and combine multiple factors for mix-and-match image generation, including sketch2color, cartoon2img, and img2gif applications. Our code/models/demo can be found at https://github.com/Yuheng-Li/MixNMatch