Corrigibility

Corrigibility
复制标题

可修正性

DOI:
--
复制
发表时间:
2015
期刊:
AI and Ethics
影响因子:
--
通讯作者:
Eliezer Yudkowsky
Eliezer Yudkowsky
中科院分区:
--
文献类型:
--
作者:
Nate Soares;Benja Fallenstein;Stuart Armstrong;Eliezer Yudkowsky

文献摘要

被引文献

相似文献

随着人工智能系统在智能和能力上的增长,它们的一些可用选项可能允许它们抵制程序员的干预。如果一个人工智能系统与它的创造者所认为的纠正性干预合作,我们就称其为“可纠正的”,尽管理性的代理人默认会抵制关闭它们或修改它们偏好的企图。我们引入了可纠正性的概念,并分析了效用函数,这些效用函数试图在关闭按钮被按下时使代理安全关闭,同时避免了阻止按钮被按下或导致按钮被按下的激励,同时确保在创建新子系统或自我修改时关闭行为的传播。虽然有些建议很有趣,但还没有一个被证明能满足我们所有的直觉需求,这使得这个简单的问题在可纠错性方面存在很大的开放性。
As artificially intelligent systems grow in intelligence and capability, some of their available options may allow them to resist intervention by their programmers. We call an AI sys-tem “corrigible” if it cooperates with what its creators regard as a corrective intervention, despite default incentives for rational agents to resist attempts to shut them down or modify their preferences. We introduce the notion of corrigibility and analyze utility functions that attempt to make an agent shut down safely if a shutdown button is pressed, while avoiding incentives to prevent the button from being pressed or cause the button to be pressed, and while ensuring propagation of the shutdown behavior as it creates new subsystems or self-modifies. While some proposals are interesting, none have yet been demonstrated to satisfy all of our intuitive desider-ata, leaving this simple problem in corrigibility wide-open.