Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation

Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation
复制标题

DOI:
10.48550/arxiv.2305.01569
复制
发表时间:
2023-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Yuval Kirstain;Adam Polyak;Uriel Singer;Shahbuland Matiana;Joe Penna;Omer Levy
Yuval Kirstain;Adam Polyak;Uriel Singer;Shahbuland Matiana;Joe Penna;Omer Levy
中科院分区:
其他
文献类型:
--
作者:
Yuval Kirstain;Adam Polyak;Uriel Singer;Shahbuland Matiana;Joe Penna;Omer Levy

文献摘要

被引文献

相似文献

从文本到图像用户收集人类偏好的大型数据集的能力通常仅限于公司,这使得公众无法访问这些数据集。为了解决这个问题,我们创建了一个Web应用程序,使文本到图像的用户生成图像,并指定他们的喜好。使用这个Web应用程序,我们构建了Pick-a-Pic,这是一个大型的开放数据集,包含文本到图像的提示和真实的用户对生成图像的偏好。我们利用这个数据集来训练一个基于CLIP的评分函数PickScore,它在预测人类偏好的任务上表现出超人的性能。然后,我们测试PickScore执行模型评估的能力,并观察到它与人类排名的相关性比其他自动评估指标更好。因此,我们建议使用PickScore来评估未来的文本到图像生成模型,并使用Pick-a-Pic提示作为比MS-COCO更相关的数据集。最后,我们展示了PickScore如何通过排名增强现有的文本到图像模型。
The ability to collect a large dataset of human preferences from text-to-image users is usually limited to companies, making such datasets inaccessible to the public. To address this issue, we create a web app that enables text-to-image users to generate images and specify their preferences. Using this web app we build Pick-a-Pic, a large, open dataset of text-to-image prompts and real users' preferences over generated images. We leverage this dataset to train a CLIP-based scoring function, PickScore, which exhibits superhuman performance on the task of predicting human preferences. Then, we test PickScore's ability to perform model evaluation and observe that it correlates better with human rankings than other automatic evaluation metrics. Therefore, we recommend using PickScore for evaluating future text-to-image generation models, and using Pick-a-Pic prompts as a more relevant dataset than MS-COCO. Finally, we demonstrate how PickScore can enhance existing text-to-image models via ranking.