Using the Veil of Ignorance to align AI systems with principles of justice.

Using the Veil of Ignorance to align AI systems with principles of justice.
复制标题

DOI:
10.1073/pnas.2213709120
复制
发表时间:
2023-05-02
影响因子:
11.1
通讯作者:
Gabriel, Iason
Gabriel, Iason
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Weidinger, Laura;Everett, Richard;Huang, Saffron;Zhu, Tina O.;Chadwick, Martin J.;Summerfield, Christopher;Fiske, Susan;Gabriel, Iason

文献摘要

参考文献

被引文献

相似文献

人工智能(AI)与社会的日益融合提出了一个关键问题:如何公平地选择原则来管理这些系统?在五项研究中,共有2,508名参与者,我们使用无知的面纱来选择原则来调整人工智能系统。与那些知道自己立场的参与者相比,面纱背后的参与者更经常选择,并在反思后认可优先考虑最糟糕情况的人工智能原则。这种模式是由对公平的更多考虑驱动的,而不是由政治取向或对风险的态度驱动的。我们的研究结果表明,无知的面纱可能是一个合适的过程,用于选择原则来管理人工智能的现实应用。哲学家约翰·罗尔斯(John Rawls)提出了无知的面纱(VoI)作为一种思想实验,以确定治理社会的公平原则。在这里,我们将VoI应用于一个重要的治理领域:人工智能(AI)。在五项激励相容的研究(N = 2508)中,包括两项预先注册的协议,参与者选择原则来从面纱后面管理人工智能(AI)助理:也就是说,不知道他们自己在群体中的相对位置。   与拥有这些信息的参与者相比,我们发现他们对指导AI助手优先考虑最差情况的原则有一致的偏好。风险态度和政治偏好都不能充分解释这些选择。相反,他们似乎是由对公平性的高度关注所驱动的:与控制条件下的参与者相比,在没有提示的情况下,在VoI背后推理的参与者更频繁地解释他们的选择是否公平。此外,我们发现VoI引发更强大偏好的能力得到了初步支持:在这里介绍的研究中,VoI增加了参与者在下一轮中继续支持他们最初选择的可能性,他们知道他们将如何受到人工智能干预的影响,并有一个自我利益的动机来改变他们的想法。这些结果出现在描述性和沉浸式游戏中。我们的研究结果表明,VoI可能是选择分配原则来管理AI的合适机制。
The growing integration of Artificial Intelligence (AI) into society raises a critical question: How can principles be fairly selected to govern these systems? Across five studies, with a total of 2,508 participants, we use the Veil of Ignorance to select principles to align AI systems. Compared to participants who know their position, participants behind the veil more frequently choose, and endorse upon reflection, principles for AI that prioritize the worst-off. This pattern is driven by increased consideration of fairness, rather than by political orientation or attitudes to risk. Our findings suggest that the Veil of Ignorance may be a suitable process for selecting principles to govern real-world applications of AI. The philosopher John Rawls proposed the Veil of Ignorance (VoI) as a thought experiment to identify fair principles for governing a society. Here, we apply the VoI to an important governance domain: artificial intelligence (AI). In five incentive-compatible studies (N = 2, 508), including two preregistered protocols, participants choose principles to govern an Artificial Intelligence (AI) assistant from behind the veil: that is, without knowledge of their own relative position in the group. Compared to participants who have this information, we find a consistent preference for a principle that instructs the AI assistant to prioritize the worst-off. Neither risk attitudes nor political preferences adequately explain these choices. Instead, they appear to be driven by elevated concerns about fairness: Without prompting, participants who reason behind the VoI more frequently explain their choice in terms of fairness, compared to those in the Control condition. Moreover, we find initial support for the ability of the VoI to elicit more robust preferences: In the studies presented here, the VoI increases the likelihood of participants continuing to endorse their initial choice in a subsequent round where they know how they will be affected by the AI intervention and have a self-interested motivation to change their mind. These results emerge in both a descriptive and an immersive game. Our findings suggest that the VoI may be a suitable mechanism for selecting distributive principles to govern AI.
DOI: 10.1007/s13347-017-0263-5
发表时间: 2018-01-01
影响因子: --
作者:
Binns, Reuben
通讯作者: Binns, Reuben
DOI: 10.1017/s0020589320000366
发表时间: 2020-10-01
影响因子: 2
作者:
Chesterman, Simon
通讯作者: Chesterman, Simon
DOI: 10.3758/s13428-021-01694-3
发表时间: 2022-08
影响因子: 5.4
作者:
Peer E;Rothschild D;Gordon A;Evernden Z;Damer E
通讯作者: Damer E
DOI: 10.1007/s11023-018-9482-5
发表时间: 2018-12-01
期刊: MINDS AND MACHINES
影响因子: 7.4
作者:
Floridi, Luciano;Cowls, Josh;Vayena, Effy
通讯作者: Vayena, Effy
DOI: 10.1111/jeea.12082
发表时间: 2014-08-01
影响因子: 3.6
作者:
Durante, Ruben;Putterman, Louis;van der Weele, Joel
通讯作者: van der Weele, Joel