CAN MACHINES LEARN MORALITY? THE DELPHI EXPERIMENT

CAN MACHINES LEARN MORALITY? THE DELPHI EXPERIMENT
复制标题

DOI:
--
复制
发表时间:
2021
影响因子:
3.5
通讯作者:
Liwei Jiang;Chandra Bhagavatula;Jenny T Liang;Jesse Dodge;Keisuke Sakaguchi;Maxwell Forbes;Jon Borchardt;Saadia Gabriel;Yulia Tsvetkov;Regina A. Rini;Yejin Choi
Liwei Jiang;Chandra Bhagavatula;Jenny T Liang;Jesse Dodge;Keisuke Sakaguchi;Maxwell Forbes;Jon Borchardt;Saadia Gabriel;Yulia Tsvetkov;Regina A. Rini;Yejin Choi
中科院分区:
法学4区
文献类型:
--
作者:
Liwei Jiang;Chandra Bhagavatula;Jenny T Liang;Jesse Dodge;Keisuke Sakaguchi;Maxwell Forbes;Jon Borchardt;Saadia Gabriel;Yulia Tsvetkov;Regina A. Rini;Yejin Choi

文献摘要

被引文献

相似文献

随着人工智能系统变得越来越强大和普遍,人们越来越担心机器的道德或缺乏道德。然而,向机器传授道德是一项艰巨的任务,因为道德仍然是人类争论最激烈的问题之一,更不用说对人工智能来说了。然而,部署在数百万用户手中的现有人工智能系统已经在做出带有道德含义的决定,这构成了一个看似不可能的挑战:教授机器道德感,而人类仍在努力应对。为了探索这一挑战,我们引入了Delphi,这是一个基于深度神经网络的实验框架,直接训练成对描述性伦理判断进行推理,例如,“帮助朋友”通常是好的,而“帮助朋友传播假新闻”不是。经验结果对机器伦理学的承诺和局限性提出了新的见解;德尔福在面对新的伦理情况时表现出强大的泛化能力,而现成的神经网络模型表现出明显的糟糕判断,包括不公正的偏见,证实了明确教授机器道德感的必要性。然而,德尔福并不完美,它容易受到普遍存在的偏见和矛盾的影响。尽管如此,我们还是展示了不完美Delphi的积极用例,包括在其他不完美的AI系统中将其用作组件模型。重要的是,我们根据突出的伦理理论来解释德尔福的可操作性,这将我们引向重要的未来研究问题。
As AI systems become increasingly powerful and pervasive, there are growing concerns about machines’ morality or a lack thereof. Yet, teaching morality to machines is a formidable task, as morality remains among the most intensely debated questions in humanity, let alone for AI. Existing AI systems deployed to millions of users, however, are already making decisions loaded with moral implications, which poses a seemingly impossible challenge: teaching machines moral sense, while humanity continues to grapple with it. To explore this challenge, we introduce Delphi, an experimental framework based on deep neural networks trained directly to reason about descriptive ethical judgments, e.g., “helping a friend” is generally good, while “helping a friend spread fake news” is not. Empirical results shed novel insights on the promises and limits of machine ethics; Delphi demonstrates strong generalization capabilities in the face of novel ethical situations, while off-the-shelf neural network models exhibit markedly poor judgment including unjust biases, confirming the need for explicitly teaching machines moral sense. Yet, Delphi is not perfect, exhibiting susceptibility to pervasive biases and inconsistencies. Despite that, we demonstrate positive use cases of imperfect Delphi, including using it as a component model within other imperfect AI systems. Importantly, we interpret the operationalization of Delphi in light of prominent ethical theories, which leads us to important future research questions.