Taxonomy of Risks posed by Language Models

Taxonomy of Risks posed by Language Models
复制标题

DOI:
10.1145/3531146.3533088
复制
发表时间:
2022-06
期刊:
Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency
影响因子:
--
通讯作者:
Laura Weidinger;J. Uesato;Maribeth Rauh;Conor Griffin;Po-Sen Huang;John F. J. Mellor;A. Glaese;Myra Cheng;Borja Balle;Atoosa Kasirzadeh;Courtney Biles;S. Brown;Zachary Kenton;W. Hawkins;T. Stepleton;Abeba Birhane;Lisa Anne Hendricks;Laura Rimell;William S. Isaac;Julia Haas;Sean Legassick;G. Irving;Iason Gabriel
Laura Weidinger;J. Uesato;Maribeth Rauh;Conor Griffin;Po-Sen Huang;John F. J. Mellor;A. Glaese;Myra Cheng;Borja Balle;Atoosa Kasirzadeh;Courtney Biles;S. Brown;Zachary Kenton;W. Hawkins;T. Stepleton;Abeba Birhane;Lisa Anne Hendricks;Laura Rimell;William S. Isaac;Julia Haas;Sean Legassick;G. Irving;Iason Gabriel
中科院分区:
其他
文献类型:
--
作者:
Laura Weidinger;J. Uesato;Maribeth Rauh;Conor Griffin;Po-Sen Huang;John F. J. Mellor;A. Glaese;Myra Cheng;Borja Balle;Atoosa Kasirzadeh;Courtney Biles;S. Brown;Zachary Kenton;W. Hawkins;T. Stepleton;Abeba Birhane;Lisa Anne Hendricks;Laura Rimell;William S. Isaac;Julia Haas;Sean Legassick;G. Irving;Iason Gabriel

文献摘要

被引文献

相似文献

大规模语言模型(LMs)的负责任创新需要对这些模型可能带来的风险有远见和深入的了解。本文发展了与lm相关的道德和社会风险的综合分类。我们根据计算机科学、语言学和社会科学的专业知识和文献,确定了21个风险。我们将这些风险归类为六个风险领域:1 .歧视、仇恨言论和排斥;信息危害,三。错误信息危害,IV.恶意使用,V.人机交互危害,VI.环境和社会经济危害。对于在lm中已经观察到的风险,讨论了导致损害的因果机制、风险的证据和减轻风险的方法。我们进一步描述和分析了尚未观察到的风险,但基于对其他语言技术的评估可以预测这些风险,并将它们置于同一分类中。我们强调,组织有责任参与我们在整个文件中讨论的缓解措施。最后,我们强调了在风险评估和缓解方面进一步研究的挑战和方向,目的是确保以负责任的方式开发语言模型。
Responsible innovation on large-scale Language Models (LMs) requires foresight into and in-depth understanding of the risks these models may pose. This paper develops a comprehensive taxonomy of ethical and social risks associated with LMs. We identify twenty-one risks, drawing on expertise and literature from computer science, linguistics, and the social sciences. We situate these risks in our taxonomy of six risk areas: I. Discrimination, Hate speech and Exclusion, II. Information Hazards, III. Misinformation Harms, IV. Malicious Uses, V. Human-Computer Interaction Harms, and VI. Environmental and Socioeconomic harms. For risks that have already been observed in LMs, the causal mechanism leading to harm, evidence of the risk, and approaches to risk mitigation are discussed. We further describe and analyse risks that have not yet been observed but are anticipated based on assessments of other language technologies, and situate these in the same taxonomy. We underscore that it is the responsibility of organizations to engage with the mitigations we discuss throughout the paper. We close by highlighting challenges and directions for further research on risk evaluation and mitigation with the goal of ensuring that language models are developed responsibly.