Limitations of Face Image Generation

Limitations of Face Image Generation
复制标题

DOI:
10.1609/aaai.v38i13.29403
复制
发表时间:
2023-09
期刊:
--
影响因子:
--
通讯作者:
Harrison Rosenberg;Shimaa Ahmed;Guruprasad V Ramesh;Ramya Korlakai Vinayak;Kassem Fawaz
Harrison Rosenberg;Shimaa Ahmed;Guruprasad V Ramesh;Ramya Korlakai Vinayak;Kassem Fawaz
中科院分区:
其他
文献类型:
--
作者:
Harrison Rosenberg;Shimaa Ahmed;Guruprasad V Ramesh;Ramya Korlakai Vinayak;Kassem Fawaz

文献摘要

相似文献

文本到图像的扩散模型由于其前所未有的图像生成能力而获得了广泛的流行。特别是,它们合成和修改人脸的能力刺激了在训练数据增强和模型性能评估中使用生成的人脸图像的研究。在本文中,我们研究了面部生成背景下生成模型的功效和缺点。利用定性和定量测量的结合,包括基于嵌入的指标和用户研究,我们提出了一个框架来审核基于一组社交属性的生成面孔的特征。我们将我们的框架应用于通过最先进的文本到图像扩散模型生成的面孔。我们确定了面部图像生成的几个局限性,包括忠实于文本提示、人口统计差异和分布变化。此外,我们提出了一个分析模型,可以深入了解训练数据选择如何影响生成模型的性能。我们的调查数据和分析代码可以在线找到:https://github.com/wi-pi/Limitations_of_Face_Generation
Text-to-image diffusion models have achieved widespread popularity due to their unprecedented image generation capability. In particular, their ability to synthesize and modify human faces has spurred research into using generated face images in both training data augmentation and model performance assessments. In this paper, we study the efficacy and shortcomings of generative models in the context of face generation. Utilizing a combination of qualitative and quantitative measures, including embedding-based metrics and user studies, we present a framework to audit the characteristics of generated faces conditioned on a set of social attributes. We applied our framework on faces generated through state-of-the-art text-to-image diffusion models. We identify several limitations of face image generation that include faithfulness to the text prompt, demographic disparities, and distributional shifts. Furthermore, we present an analytical model that provides insights into how training data selection contributes to the performance of generative models. Our survey data and analytics code can be found online at https://github.com/wi-pi/Limitations_of_Face_Generation