LD-ZNet: A Latent Diffusion Approach for Text-Based Image Segmentation

LD-ZNet: A Latent Diffusion Approach for Text-Based Image Segmentation
复制标题

DOI:
10.1109/iccv51070.2023.00384
复制
发表时间:
2023-03
期刊:
2023 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
通讯作者:
K. Pnvr;Bharat Singh;P. Ghosh;Behjat Siddiquie;David Jacobs
K. Pnvr;Bharat Singh;P. Ghosh;Behjat Siddiquie;David Jacobs
中科院分区:
其他
文献类型:
--
作者:
K. Pnvr;Bharat Singh;P. Ghosh;Behjat Siddiquie;David Jacobs

文献摘要

被引文献

相似文献

大规模的预训练任务,如图像分类、字幕或自监督技术,不会激励学习对象的语义边界。然而,最近使用基于文本的潜在扩散技术建立的生成基础模型可以学习语义边界。这是因为他们必须根据文本描述综合图像中所有物体的复杂细节。因此,我们提出了一种使用在互联网规模数据集上训练的潜在扩散模型(ldm)分割真实图像和人工智能生成图像的技术。首先,我们证明了与RGB图像或CLIP编码等其他特征表示相比,ldm的潜在空间(z空间)是一种更好的输入表示,用于基于文本的图像分割。通过在潜在z空间上训练分割模型,在不同形式的艺术、漫画、插图和照片等多个领域中创建压缩表示,我们还能够弥合真实图像和人工智能生成图像之间的领域差距。我们证明了ldm的内部特征包含了丰富的语义信息,并提出了一种以LD-ZNet形式的技术来进一步提高基于文本的分割性能。总的来说,我们在自然图像的文本到图像分割方面比标准基线提高了6%。对于人工智能生成的图像,与最先进的技术相比,我们显示了近20%的改进。该项目可在https://koutilya-pnvr.github.io/LD-ZNet/上获得。
Large-scale pre-training tasks like image classification, captioning, or self-supervised techniques do not incentivize learning the semantic boundaries of objects. However, recent generative foundation models built using text-based latent diffusion techniques may learn semantic boundaries. This is because they have to synthesize intricate details about all objects in an image based on a text description. Therefore, we present a technique for segmenting real and AI-generated images using latent diffusion models (LDMs) trained on internet-scale datasets. First, we show that the latent space of LDMs (z-space) is a better input representation compared to other feature representations like RGB images or CLIP encodings for text-based image segmentation. By training the segmentation models on the latent z-space, which creates a compressed representation across several domains like different forms of art, cartoons, illustrations, and photographs, we are also able to bridge the domain gap between real and AI-generated images. We show that the internal features of LDMs contain rich semantic information and present a technique in the form of LD-ZNet to further boost the performance of text-based segmentation. Overall, we show up to 6% improvement over standard baselines for text-to-image segmentation on natural images. For AI-generated imagery, we show close to 20% improvement compared to state-of-the-art techniques. The project is available at https://koutilya-pnvr.github.io/LD-ZNet/.