On Second Thought, Let’s Not Think Step by Step! Bias and Toxicity in Zero-Shot Reasoning

On Second Thought, Let’s Not Think Step by Step! Bias and Toxicity in Zero-Shot Reasoning
复制标题

DOI:
10.48550/arxiv.2212.08061
复制
发表时间:
2022-12
期刊:
ArXiv
影响因子:
--
通讯作者:
Omar Shaikh;Hongxin Zhang;William B. Held;Michael Bernstein;Diyi Yang
Omar Shaikh;Hongxin Zhang;William B. Held;Michael Bernstein;Diyi Yang
中科院分区:
其他
文献类型:
--
作者:
Omar Shaikh;Hongxin Zhang;William B. Held;Michael Bernstein;Diyi Yang

文献摘要

相似文献

生成思想链(CoT)已被证明可以在各种NLP任务上持续提高大型语言模型(LLM)的性能。然而,以前的工作主要集中在逻辑推理任务(例如算术,常识问答);目前还不清楚是否有更多样化的推理类型,特别是在社会情境下的改进。具体而言,我们进行了控制评估零杆CoT在两个社会敏感领域:有害的问题和刻板印象基准。我们发现,零杆CoT推理在敏感领域显着增加了模型的可能性,产生有害或不受欢迎的输出,在不同的提示格式和模型变体的趋势举行。此外,我们还表明,有害的CoT随着模型大小的增加而增加,但随着指令遵循的改善而减少。我们的工作表明,零镜头CoT应谨慎使用社会重要的任务,特别是当涉及边缘化群体或敏感话题。
Generating a Chain of Thought (CoT) has been shown to consistently improve large language model (LLM) performance on a wide range of NLP tasks. However, prior work has mainly focused on logical reasoning tasks (e.g. arithmetic, commonsense QA); it remains unclear whether improvements hold for more diverse types of reasoning, especially in socially situated contexts. Concretely, we perform a controlled evaluation of zero-shot CoT across two socially sensitive domains: harmful questions and stereotype benchmarks. We find that zero-shot CoT reasoning in sensitive domains significantly increases a model’s likelihood to produce harmful or undesirable output, with trends holding across different prompt formats and model variants. Furthermore, we show that harmful CoTs increase with model size, but decrease with improved instruction following. Our work suggests that zero-shot CoT should be used with caution on socially important tasks, especially when marginalized groups or sensitive topics are involved.