Ethical and social risks of harm from Language Models

Ethical and social risks of harm from Language Models
复制标题

DOI:
--
复制
发表时间:
2021-12
期刊:
ArXiv
影响因子:
--
通讯作者:
Laura Weidinger;John F. J. Mellor;Maribeth Rauh;Conor Griffin;J. Uesato;Po-Sen Huang;Myra Cheng;Mia Glaese;Borja Balle;Atoosa Kasirzadeh;Zachary Kenton;S. Brown;W. Hawkins;T. Stepleton;Courtney Biles;Abeba Birhane;Julia Haas;Laura Rimell;Lisa Anne Hendricks;William S. Isaac;Sean Legassick;G. Irving;Iason Gabriel
Laura Weidinger;John F. J. Mellor;Maribeth Rauh;Conor Griffin;J. Uesato;Po-Sen Huang;Myra Cheng;Mia Glaese;Borja Balle;Atoosa Kasirzadeh;Zachary Kenton;S. Brown;W. Hawkins;T. Stepleton;Courtney Biles;Abeba Birhane;Julia Haas;Laura Rimell;Lisa Anne Hendricks;William S. Isaac;Sean Legassick;G. Irving;Iason Gabriel
中科院分区:
其他
文献类型:
--
作者:
Laura Weidinger;John F. J. Mellor;Maribeth Rauh;Conor Griffin;J. Uesato;Po-Sen Huang;Myra Cheng;Mia Glaese;Borja Balle;Atoosa Kasirzadeh;Zachary Kenton;S. Brown;W. Hawkins;T. Stepleton;Courtney Biles;Abeba Birhane;Julia Haas;Laura Rimell;Lisa Anne Hendricks;William S. Isaac;Sean Legassick;G. Irving;Iason Gabriel

文献摘要

被引文献

相似文献

本文旨在帮助构建与大规模语言模型(LM)相关的风险格局。为了促进负责任的创新,需要深入了解这些模式带来的潜在风险。广泛的既定和预期的风险进行了详细分析,借鉴多学科的专业知识和文献,从计算机科学,语言学和社会科学。我们概括了六个具体的风险领域:一。歧视、排斥和毒性,II。信息危害,三。错误信息危害,V.恶意使用,V.人机交互危害,VI.自动化、访问和环境危害。第一个领域涉及陈规定型观念的延续、不公平的歧视、排斥性规范、有毒语言和社会群体对地方妇女的低绩效。第二个重点是私人数据泄露或LM正确推断敏感信息的风险。第三类涉及不良、虚假或误导性信息(包括敏感领域的信息)所产生的风险,以及对共享信息的信任度下降等连锁风险。第四个考虑的是试图使用地雷造成伤害的行为者的风险。第五个重点是用于支持与人类用户交互的会话代理的LLM特有的风险,包括不安全的使用,操纵或欺骗。第六章讨论了环境危害、工作自动化和其他挑战的风险,这些挑战可能对不同的社会群体或社区产生不同的影响。我们总共深入审查了21项风险。我们讨论了不同风险的起源点,并指出潜在的缓解方法。最后,我们讨论了实施缓解措施的组织责任,以及合作和参与的作用。我们强调了进一步研究的方向,特别是在扩大评估和评估LM中概述的风险的工具包。
This paper aims to help structure the risk landscape associated with large-scale Language Models (LMs). In order to foster advances in responsible innovation, an in-depth understanding of the potential risks posed by these models is needed. A wide range of established and anticipated risks are analysed in detail, drawing on multidisciplinary expertise and literature from computer science, linguistics, and social sciences. We outline six specific risk areas: I. Discrimination, Exclusion and Toxicity, II. Information Hazards, III. Misinformation Harms, V. Malicious Uses, V. Human-Computer Interaction Harms, VI. Automation, Access, and Environmental Harms. The first area concerns the perpetuation of stereotypes, unfair discrimination, exclusionary norms, toxic language, and lower performance by social group for LMs. The second focuses on risks from private data leaks or LMs correctly inferring sensitive information. The third addresses risks arising from poor, false or misleading information including in sensitive domains, and knock-on risks such as the erosion of trust in shared information. The fourth considers risks from actors who try to use LMs to cause harm. The fifth focuses on risks specific to LLMs used to underpin conversational agents that interact with human users, including unsafe use, manipulation or deception. The sixth discusses the risk of environmental harm, job automation, and other challenges that may have a disparate effect on different social groups or communities. In total, we review 21 risks in-depth. We discuss the points of origin of different risks and point to potential mitigation approaches. Lastly, we discuss organisational responsibilities in implementing mitigations, and the role of collaboration and participation. We highlight directions for further research, particularly on expanding the toolkit for assessing and evaluating the outlined risks in LMs.