Protecting Against Malicious Use of Image Diffusion Models
Protecting Against Malicious Use of Image Diffusion Models
批准号:
2737559
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2022
资助国家:
英国
项目状态:
未结题
起止时间:
2022 至 --
中文摘要
点击翻译按钮获取中文摘要
英文摘要
The rapid advancement in the capabilities of generative AI has been accompanied by an increase in scope for their misuse. In particular, text-to-image diffusion models such as DALL-E 2 have reacheda point at which high quality images can be generated with ease by a broad spectrum of users, not all of whom may have good intent. The aim of this PhD project is to investigate how to protect againstmalicious use of this technology. The goal of a text-to-image diffusion model is to generate realistic images based on a text prompt. These models work by using a forward process which sequentially adds noise to an input image and then learns a mapping to reverse this process. Then, starting with noise, we can use the trained model to denoise the image conditioned on the input text and end up with a output image that appears to come from the same distribution as the input images. Instead of retraining or changing how these models work, we can modify parameters in existing models to change their behaviour to be more desirable. One example use case for this methodology is editing implicit biases found in these models. Implicit biases in the training data for text-to-image diffusion models, can lead to perpetuating social and cultural biases when generating images. As an example, asking a diffusion model for an image of a cow will, with high probability, return an image of a cow in a field even although the environment was never specified. Gender biases are also present in these models which is noticeable when generating pictures of people in certain professions. Since retraining the model to avoid these biases is expensive and time consuming, methods such as TIME look to edit the weights of the model after training in order to reduce the likelihood of a chosen bias occurring. There is still a large amount of analysis to be done into the methodology. Some examples include: How is model performance impacted after editing facts? How can we reduce a broader range of biases in one application of the TIME method? Image manipulation using image diffusion models is another emerging issue. The availability and ease at which these tools can be used allow any image to be edited freely, with little in the way of safeguards to prevent malicious intent. In-painting is the process of taking an existing image and using the model to only generate specific areas of the image. One example of how this can be misused is by editing the background of an image of a person to make it appear as if they were somewhere else. Methodology has been proposed to create protections for images against this process, by adding specific perturbations to the image that cause diffusion models to struggle to generate what is prompted. This project would include an analysis into these techniques and other potential prevention strategies.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金