StyleT2I: Toward Compositional and High-Fidelity Text-to-Image Synthesis

StyleT2I: Toward Compositional and High-Fidelity Text-to-Image Synthesis
复制标题

DOI:
10.1109/cvpr52688.2022.01766
复制
发表时间:
2022-03
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Zhiheng Li;Martin Renqiang Min;K. Li;Chenliang Xu
Zhiheng Li;Martin Renqiang Min;K. Li;Chenliang Xu
中科院分区:
其他
文献类型:
--
作者:
Zhiheng Li;Martin Renqiang Min;K. Li;Chenliang Xu

文献摘要

相似文献

虽然已经取得了进展,文本到图像的合成,以前的方法没有推广到看不见的或代表性不足的属性组成的输入文本。缺乏组合性可能对鲁棒性和公平性产生严重影响,例如,无法合成代表性不足的人口群体的面部图像。在本文中,我们引入了一个新的框架,StyleT 2 I,以提高文本到图像合成的组合性。具体来说,我们提出了一个CLIP引导的对比损失,以更好地区分不同的句子之间的不同成分。为了进一步提高组合性,我们设计了一种新的语义匹配损失和空间约束,以确定属性的潜在方向,为预期的空间区域操作,从而更好地解开属性的潜在表示。在此基础上,提出了组合属性调整算法,对图像的潜在编码进行调整,提高了图像合成的组合性。此外,我们利用$l_{2}$-norm正则化识别的潜在方向(范数惩罚),以达到一个很好的平衡之间的图像-文本对齐和图像保真度。在实验中,我们设计了一个新的数据集分裂和评价指标来评估文本到图像合成模型的组合性。实验结果表明,StyleT 2 I在输入文本与合成图像的一致性方面优于以往的方法,并实现了更高的保真度。
Although progress has been made for text-to-image synthesis, previous methods fall short of generalizing to unseen or underrepresented attribute compositions in the input text. Lacking compositionality could have severe implications for robustness and fairness, e.g., inability to synthesize the face images of underrepresented demographic groups. In this paper, we introduce a new framework, StyleT2I, to improve the compositionality of text-to-image synthesis. Specifically, we propose a CLIP-guided Contrastive Loss to better distinguish different compositions among different sentences. To further improve the compositionality, we design a novel Semantic Matching Loss and a Spatial Constraint to identify attributes' latent directions for intended spatial region manipulations, leading to better disentangled latent representations of attributes. Based on the identified latent directions of attributes, we propose Compositional Attribute Adjustment to adjust the latent code, resulting in better compositionality of image synthesis. In addition, we leverage the $l_{2}$-norm regularization of identified latent directions (norm penalty) to strike a nice balance between image-text alignment and image fidelity. In the experiments, we devise a new dataset split and an evaluation metric to evaluate the compositionality of text-to-image synthesis models. The results show that StyleT2I outperforms previous approaches in terms of the consistency between the input text and synthesized images and achieves higher fidelity.