Text-based Image Editing for Food Images with CLIP
Text-based Image Editing for Food Images with CLIP
复制标题
DOI:
10.1145/3552484.3555751
复制
发表时间:
2022-10
期刊:
影响因子:
--
通讯作者:
Kohei Yamamoto;Keiji Yanai
中科院分区:
文献类型:
--
作者:
Kohei Yamamoto;Keiji Yanai
Recently, the large-scale language-image pre-trained model, such as CLIP, has drawn much attention due to its remarkable ability for various tasks, including classification and image synthesis. The combination of CLIP and GAN can be used for text-based image manipulation and text-based image synthesis.Several models of a combination of CLIP and GAN have been proposed so far. However, their effectiveness in the food image domain has not been examined comprehensively yet. In this paper, we reported the results of the experiments on text-based food image manipulation using VQGAN-CLIP and discussed the possibility of food image manipulation by texts.