Aligning AI With Shared Human Values

Aligning AI With Shared Human Values
复制标题

DOI:
--
复制
发表时间:
2020-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Dan Hendrycks;Collin Burns;Steven Basart;Andrew Critch;J. Li;D. Song;J. Steinhardt
Dan Hendrycks;Collin Burns;Steven Basart;Andrew Critch;J. Li;D. Song;J. Steinhardt
中科院分区:
其他
文献类型:
--
作者:
Dan Hendrycks;Collin Burns;Steven Basart;Andrew Critch;J. Li;D. Song;J. Steinhardt

文献摘要

被引文献

相似文献

我们展示了如何评估语言模型对道德基本概念的了解。我们介绍了伦理数据集,这是一个新的基准,涵盖了正义、福祉、责任、美德和常识道德等概念。模型预测了对不同文本情景的普遍道德判断。这需要将物理和社会世界的知识与价值判断联系起来,这种能力可能使我们能够引导聊天机器人的输出,或者最终使开放式强化学习代理正规化。使用伦理学数据集,我们发现目前的语言模型对基本的伦理学知识有一个有希望但不完整的理解。我们的工作表明,今天可以在机器伦理方面取得进展,这为人工智能提供了一个与人类价值观保持一致的踏脚石。
We show how to assess a language model's knowledge of basic concepts of morality. We introduce the ETHICS dataset, a new benchmark that spans concepts in justice, well-being, duties, virtues, and commonsense morality. Models predict widespread moral judgments about diverse text scenarios. This requires connecting physical and social world knowledge to value judgements, a capability that may enable us to steer chatbot outputs or eventually regularize open-ended reinforcement learning agents. With the ETHICS dataset, we find that current language models have a promising but incomplete understanding of basic ethical knowledge. Our work shows that progress can be made on machine ethics today, and it provides a steppingstone toward AI that is aligned with human values.