Language Generation Models Can Cause Harm: So What Can We Do About It? An Actionable Survey

Language Generation Models Can Cause Harm: So What Can We Do About It? An Actionable Survey
复制标题

DOI:
10.48550/arxiv.2210.07700
复制
发表时间:
2022-10
期刊:
--
影响因子:
--
通讯作者:
Sachin Kumar;Vidhisha Balachandran;Lucille Njoo;Antonios Anastasopoulos;Yulia Tsvetkov
Sachin Kumar;Vidhisha Balachandran;Lucille Njoo;Antonios Anastasopoulos;Yulia Tsvetkov
中科院分区:
其他
文献类型:
--
作者:
Sachin Kumar;Vidhisha Balachandran;Lucille Njoo;Antonios Anastasopoulos;Yulia Tsvetkov

文献摘要

相似文献

大型语言模型生成类人文本能力的最新进展导致它们在面向用户的环境中越来越多地被采用。与此同时,这些改进也引发了一场激烈的讨论,围绕着它们所带来的社会危害的风险,无论是无意的还是恶意的。一些研究已经探索了这些危害,并呼吁通过开发更安全,更公平的模型来减轻这些危害。除了列举危害的风险之外,这项工作还提供了一个实用方法的调查,以解决语言生成模型的潜在威胁和社会危害。我们借鉴了几个以前的作品的语言模型风险的分类,提出了一个结构化的概述策略,用于检测和改善不同类型的风险/危害的语言生成器。本调查旨在为LM研究人员和实践者提供实用指南,并解释不同策略的动机,其局限性以及未来研究的开放问题。
Recent advances in the capacity of large language models to generate human-like text have resulted in their increased adoption in user-facing settings. In parallel, these improvements have prompted a heated discourse around the risks of societal harms they introduce, whether inadvertent or malicious. Several studies have explored these harms and called for their mitigation via development of safer, fairer models. Going beyond enumerating the risks of harms, this work provides a survey of practical methods for addressing potential threats and societal harms from language generation models. We draw on several prior works’ taxonomies of language model risks to present a structured overview of strategies for detecting and ameliorating different kinds of risks/harms of language generators. Bridging diverse strands of research, this survey aims to serve as a practical guide for both LM researchers and practitioners, with explanations of different strategies’ motivations, their limitations, and open problems for future research.