Seeing Seeds Beyond Weeds: Green Teaming Generative AI for Beneficial Uses

Seeing Seeds Beyond Weeds: Green Teaming Generative AI for Beneficial Uses
复制标题

DOI:
10.48550/arxiv.2306.03097
复制
发表时间:
2023-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Logan Stapleton;Jordan Taylor;Sarah Fox;Tongshuang Sherry Wu;Haiyi Zhu
Logan Stapleton;Jordan Taylor;Sarah Fox;Tongshuang Sherry Wu;Haiyi Zhu
中科院分区:
其他
文献类型:
--
作者:
Logan Stapleton;Jordan Taylor;Sarah Fox;Tongshuang Sherry Wu;Haiyi Zhu

文献摘要

被引文献

相似文献

GPT和DALL-E等大型生成式AI模型(GM)经过训练,可以生成用于一般、广泛用途的内容。GM内容过滤器一般用于过滤掉在许多情况下具有伤害风险的内容,例如,仇恨言论然而,被禁止的内容并不总是有害的--在某些情况下,生成被禁止的内容可能是有益的。因此,当GM过滤掉内容时,他们会将有益的用例沿着有害的用例一起排除。排除哪些用例反映了GM内容过滤中嵌入的值。最近关于红色团队的工作提出了绕过转基因内容过滤器以生成有害内容的方法。我们创造了绿色团队这个术语来描述绕过GM内容过滤器以设计有益用例的方法。我们通过以下方式展示绿色团队合作:1)使用ChatGPT作为虚拟患者来模拟一个有自杀想法的人,进行自杀支持培训; 2)使用Codex故意生成错误的解决方案,以培训学生调试; 3)使用Midjourney检查Instagram页面,以生成反LGBTQ+政客的图像。最后,我们讨论了我们的用例如何将绿色团队作为一种实用的设计方法和一种批判模式来展示,它质疑并颠覆了当前对生成式人工智能的危害和价值的理解。
Large generative AI models (GMs) like GPT and DALL-E are trained to generate content for general, wide-ranging purposes. GM content filters are generalized to filter out content which has a risk of harm in many cases, e.g., hate speech. However, prohibited content is not always harmful -- there are instances where generating prohibited content can be beneficial. So, when GMs filter out content, they preclude beneficial use cases along with harmful ones. Which use cases are precluded reflects the values embedded in GM content filtering. Recent work on red teaming proposes methods to bypass GM content filters to generate harmful content. We coin the term green teaming to describe methods of bypassing GM content filters to design for beneficial use cases. We showcase green teaming by: 1) Using ChatGPT as a virtual patient to simulate a person experiencing suicidal ideation, for suicide support training; 2) Using Codex to intentionally generate buggy solutions to train students on debugging; and 3) Examining an Instagram page using Midjourney to generate images of anti-LGBTQ+ politicians in drag. Finally, we discuss how our use cases demonstrate green teaming as both a practical design method and a mode of critique, which problematizes and subverts current understandings of harms and values in generative AI.