Corrigibility
Corrigibility
复制标题
可修正性
DOI:
--
复制
发表时间:
2015
期刊:
影响因子:
--
通讯作者:
Eliezer Yudkowsky
中科院分区:
文献类型:
--
作者:
Nate Soares;Benja Fallenstein;Stuart Armstrong;Eliezer Yudkowsky
As artificially intelligent systems grow in intelligence and capability, some of their available options may allow them to resist intervention by their programmers. We call an AI sys-tem “corrigible” if it cooperates with what its creators regard as a corrective intervention, despite default incentives for rational agents to resist attempts to shut them down or modify their preferences. We introduce the notion of corrigibility and analyze utility functions that attempt to make an agent shut down safely if a shutdown button is pressed, while avoiding incentives to prevent the button from being pressed or cause the button to be pressed, and while ensuring propagation of the shutdown behavior as it creates new subsystems or self-modifies. While some proposals are interesting, none have yet been demonstrated to satisfy all of our intuitive desider-ata, leaving this simple problem in corrigibility wide-open.