Text-based Image Editing for Food Images with CLIP

Text-based Image Editing for Food Images with CLIP
复制标题

DOI:
10.1145/3552484.3555751
复制
发表时间:
2022-10
期刊:
Proceedings of the 7th International Workshop on Multimedia Assisted Dietary Management
影响因子:
--
通讯作者:
Kohei Yamamoto;Keiji Yanai
Kohei Yamamoto;Keiji Yanai
中科院分区:
其他
文献类型:
--
作者:
Kohei Yamamoto;Keiji Yanai

文献摘要

相似文献

近年来,以CLIP为代表的大规模语言图像预训练模型因其在分类、图像合成等方面的卓越性能而备受关注。CLIP和GaN的结合可以用于基于文本的图像处理和基于文本的图像合成。目前已经提出了几种CLIP和GaN的组合模型。然而,它们在食品形象领域的有效性还没有得到全面的检验。本文报道了利用VQGAN-CLIP进行基于文本的食物图像操纵的实验结果,并讨论了文本操纵食物图像的可能性。
Recently, the large-scale language-image pre-trained model, such as CLIP, has drawn much attention due to its remarkable ability for various tasks, including classification and image synthesis. The combination of CLIP and GAN can be used for text-based image manipulation and text-based image synthesis.Several models of a combination of CLIP and GAN have been proposed so far. However, their effectiveness in the food image domain has not been examined comprehensively yet. In this paper, we reported the results of the experiments on text-based food image manipulation using VQGAN-CLIP and discussed the possibility of food image manipulation by texts.