Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
复制标题

DOI:
10.48550/arxiv.2209.07858
复制
发表时间:
2022-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Deep Ganguli;Liane Lovitt;John Kernion;Amanda Askell;Yuntao Bai;Saurav Kadavath;Benjamin Mann;Ethan Perez;Nicholas Schiefer;Kamal Ndousse;Andy Jones;Sam Bowman;Anna Chen;Tom Conerly;Nova Dassarma;Dawn Drain;Nelson Elhage;S. El-Showk;Stanislav Fort;Z. Dodds;T. Henighan;Danny Hernandez;Tristan Hume;Josh Jacobson;Scott Johnston;Shauna Kravec;Catherine Olsson;Sam Ringer;Eli Tran-Johnson;Dario Amodei;Tom B. Brown;Nicholas Joseph;Sam McCandlish;C. Olah;Jared Kaplan;Jack Clark
Deep Ganguli;Liane Lovitt;John Kernion;Amanda Askell;Yuntao Bai;Saurav Kadavath;Benjamin Mann;Ethan Perez;Nicholas Schiefer;Kamal Ndousse;Andy Jones;Sam Bowman;Anna Chen;Tom Conerly;Nova Dassarma;Dawn Drain;Nelson Elhage;S. El-Showk;Stanislav Fort;Z. Dodds;T. Henighan;Danny Hernandez;Tristan Hume;Josh Jacobson;Scott Johnston;Shauna Kravec;Catherine Olsson;Sam Ringer;Eli Tran-Johnson;Dario Amodei;Tom B. Brown;Nicholas Joseph;Sam McCandlish;C. Olah;Jared Kaplan;Jack Clark
中科院分区:
其他
文献类型:
--
作者:
Deep Ganguli;Liane Lovitt;John Kernion;Amanda Askell;Yuntao Bai;Saurav Kadavath;Benjamin Mann;Ethan Perez;Nicholas Schiefer;Kamal Ndousse;Andy Jones;Sam Bowman;Anna Chen;Tom Conerly;Nova Dassarma;Dawn Drain;Nelson Elhage;S. El-Showk;Stanislav Fort;Z. Dodds;T. Henighan;Danny Hernandez;Tristan Hume;Josh Jacobson;Scott Johnston;Shauna Kravec;Catherine Olsson;Sam Ringer;Eli Tran-Johnson;Dario Amodei;Tom B. Brown;Nicholas Joseph;Sam McCandlish;C. Olah;Jared Kaplan;Jack Clark

文献摘要

被引文献

相似文献

我们描述了我们对红队语言模型的早期努力,以便同时发现,测量和尝试减少其潜在的有害输出。我们做出了三个主要贡献。首先,我们研究了3种模型大小(2.7B,13 B和52 B参数)和4种模型类型的红色团队的缩放行为:普通语言模型(LM);提示有用,诚实和无害的LM;具有拒绝采样的LM;以及使用来自人类反馈的强化学习(RLHF)训练为有用和无害的模型。我们发现,RLHF模型越来越难以红队,因为他们的规模,我们发现一个平坦的趋势与规模的其他模型类型。其次,我们发布了38,961次红队攻击的数据集,供其他人分析和学习。我们提供了自己的数据分析,并发现了各种有害的输出,从攻击性语言到更微妙的有害的非暴力不道德输出。第三,我们详尽地描述了我们的指示,流程,统计方法,以及关于红色团队的不确定性。我们希望这种透明度能够加速我们作为一个社区共同工作的能力,以便为如何红团队语言模型开发共享的规范,实践和技术标准。
We describe our early efforts to red team language models in order to simultaneously discover, measure, and attempt to reduce their potentially harmful outputs. We make three main contributions. First, we investigate scaling behaviors for red teaming across 3 model sizes (2.7B, 13B, and 52B parameters) and 4 model types: a plain language model (LM); an LM prompted to be helpful, honest, and harmless; an LM with rejection sampling; and a model trained to be helpful and harmless using reinforcement learning from human feedback (RLHF). We find that the RLHF models are increasingly difficult to red team as they scale, and we find a flat trend with scale for the other model types. Second, we release our dataset of 38,961 red team attacks for others to analyze and learn from. We provide our own analysis of the data and find a variety of harmful outputs, which range from offensive language to more subtly harmful non-violent unethical outputs. Third, we exhaustively describe our instructions, processes, statistical methodologies, and uncertainty about red teaming. We hope that this transparency accelerates our ability to work together as a community in order to develop shared norms, practices, and technical standards for how to red team language models.