Protecting Against Malicious Use of Image Diffusion Models
Protecting Against Malicious Use of Image Diffusion Models
批准号:
2737559
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2022
资助国家:
英国
项目状态:
未结题
起止时间:
2022 至 --
中文摘要
生成性人工智能能力的快速进步伴随着它们被滥用的范围的增加。特别是,Dall-E2等文本到图像扩散模型已经达到了一个点,在这个点上,广大用户可以轻松地生成高质量的图像,而不是所有用户都有良好的意图。这个博士项目的目的是研究如何防止恶意使用这项技术。文本到图像扩散模型的目标是基于文本提示生成逼真的图像。这些模型的工作原理是使用正向过程,该过程按顺序将噪声添加到输入图像,然后学习映射以逆转这一过程。然后,从噪声开始,我们可以使用训练好的模型来对以输入文本为条件的图像进行去噪,最终得到的输出图像似乎来自与输入图像相同的分布。我们不需要再培训或改变这些模型的工作方式,而是可以修改现有模型中的参数,以改变它们的行为,使其变得更令人满意。此方法的一个示例用例是编辑在这些模型中发现的隐式偏差。文本到图像扩散模型的训练数据中的隐含偏差可能会导致在生成图像时长期存在社会和文化偏差。例如,向扩散模型请求牛的图像的可能性很高,即使从未指定环境,也会返回田里的牛的图像。性别偏见也出现在这些模特身上,在制作某些职业的人的照片时,这一点很明显。由于重新训练模型以避免这些偏差是昂贵且耗时的,诸如时间的方法在训练之后寻求编辑模型的权重,以减少所选择的偏差发生的可能性。在方法论方面仍有大量的分析工作要做。一些例子包括:编辑事实后对模型性能有何影响?我们如何在一次时间方法的应用中减少更大范围的偏差?使用图像扩散模型的图像处理是另一个新出现的问题。这些工具的可用性和易用性允许对任何图像进行自由编辑,几乎没有采取任何安全措施来防止恶意意图。内绘制是获取现有图像并使用模型仅生成图像的特定区域的过程。这可能被滥用的一个例子是,通过编辑一个人的图像的背景,使其看起来就像他们在其他地方一样。已经提出了一种方法来保护图像免受这一过程的影响,方法是向图像添加特定的扰动,使扩散模型难以生成提示的内容。该项目将包括对这些技术和其他潜在预防战略的分析。
英文摘要
The rapid advancement in the capabilities of generative AI has been accompanied by an increase in scope for their misuse. In particular, text-to-image diffusion models such as DALL-E 2 have reacheda point at which high quality images can be generated with ease by a broad spectrum of users, not all of whom may have good intent. The aim of this PhD project is to investigate how to protect againstmalicious use of this technology. The goal of a text-to-image diffusion model is to generate realistic images based on a text prompt. These models work by using a forward process which sequentially adds noise to an input image and then learns a mapping to reverse this process. Then, starting with noise, we can use the trained model to denoise the image conditioned on the input text and end up with a output image that appears to come from the same distribution as the input images. Instead of retraining or changing how these models work, we can modify parameters in existing models to change their behaviour to be more desirable. One example use case for this methodology is editing implicit biases found in these models. Implicit biases in the training data for text-to-image diffusion models, can lead to perpetuating social and cultural biases when generating images. As an example, asking a diffusion model for an image of a cow will, with high probability, return an image of a cow in a field even although the environment was never specified. Gender biases are also present in these models which is noticeable when generating pictures of people in certain professions. Since retraining the model to avoid these biases is expensive and time consuming, methods such as TIME look to edit the weights of the model after training in order to reduce the likelihood of a chosen bias occurring. There is still a large amount of analysis to be done into the methodology. Some examples include: How is model performance impacted after editing facts? How can we reduce a broader range of biases in one application of the TIME method? Image manipulation using image diffusion models is another emerging issue. The availability and ease at which these tools can be used allow any image to be edited freely, with little in the way of safeguards to prevent malicious intent. In-painting is the process of taking an existing image and using the model to only generate specific areas of the image. One example of how this can be misused is by editing the background of an image of a person to make it appear as if they were somewhere else. Methodology has been proposed to create protections for images against this process, by adding specific perturbations to the image that cause diffusion models to struggle to generate what is prompted. This project would include an analysis into these techniques and other potential prevention strategies.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金