Superintelligence Does Not Imply Benevolence

Superintelligence Does Not Imply Benevolence
复制标题

超级智能并不意味着仁慈

DOI:
--
复制
发表时间:
2013
期刊:
影响因子:
--
通讯作者:
Carl Shulman
Carl Shulman
中科院分区:
--
文献类型:
--
作者:
Joshua Fox;Carl Shulman

文献摘要

参考文献

被引文献

相似文献

随着机器变得能够进行更加自主和智能的行为,它们也会表现出更多道德上可取的行为吗?地球的历史往往表明,智力、知识和理性的增加将带来更多的合作和仁慈的行为。具有复杂神经系统的动物会追踪并惩罚剥削行为,同时奖励合作。与其他动物相比,人类形成了复杂的规范和规模惊人的社会群体。即使在人类经验中,随着时间的推移,知识的积累也与暴力发生率的降低(Pinker 2007)和合作范围的扩大(Wright 2001)相关,从部落到部落到城邦到国家再到跨国组织。人们可能会从这一趋势中得出结论,并认为随着机器接近并超越人类的认知能力,福克斯、约书亚和卡尔·舒尔曼。 2010.“超级智能并不意味着仁慈。”载于 ECAP10:第八届欧洲计算与哲学会议,由 Klaus Mainzer 编辑。慕尼黑:博士小屋出版社。此版本包含细微更改。道德行为也会随之提高。我们认为,这种图景忽略了两种道德概念之间的关键区别,以及从增加智力到更加道德行为的途径之间的相关区别。一种观念将道德视为具有不同目标的实体之间合作的系统,并且有能力影响彼此对这些目标的追求。诸如互惠利他主义(Trivers 1971)之类的做法可以帮助伴侣提高各自的生殖适应性。在合作观念中,执行道德行为或安排自己这样做的原因(Gauthier 1986)是为了推进自己的目标。另一种价值论观念认为,道德要求修正我们的最终目标。这一概念对于治疗无助的动物(例如非人类动物)尤其重要。合作道德理论,例如高蒂尔(Gauthier,1986),通常只能通过与利他的强大代理人的合作来获得无助者的道德地位。然后我们可以评估从智力到道德行为的替代路径。首先,具有更强工具理性的机器可以更好地设计和实施合作实践。因此,霍尔(Hall,2007)认为,智能机器将超越人类,至少与强大的同行合作。其次,根据康德的道德观,人们可能会认为,随着智能机器扩展其知识和能力,它们会直接受到激励而修改自己的偏好,使其更加道德(Chalmers 2010)。我们考虑康德观点的一个特定的反例。使用智能定义为在广泛的环境中实现目标的能力(Legg 2008),我们讨论了 AIXI 形式主义,它将所罗门诺夫归纳法与贝叶斯决策理论相结合,以优化未知的奖励函数(Hutter 2005)。 AIXI 虽然在物理上无法实现,但它是一种紧凑指定的超级智能,在最大化任意目标方面可证明是最佳的,但“没有空间”进行康德式修正。相反,它在大多数情况下会保留任意值(Omohundro 2008)。因此,我们有理由认为,出于工具性原因,不同的智能机器会集中表现出与足够强大的合作伙伴合作的“动力”,即使这不是专门设计的。然而,我们有理由对未经精心设计的利他性智能机器的最终目的感到悲观,因此应该努力避免此类系统相对于人类而言非常强大的情况(Yudkowsky 2008)。约书亚·福克斯,卡尔·舒尔曼
As machines become capable of more autonomous and intelligent behavior, will they also display more morally desirable behavior? Earth’s history tends to suggest that increasing intelligence, knowledge, and rationality will result in more cooperative and benevolent behavior. Animals with sophisticated nervous systems track and punish exploitative behavior, while rewarding cooperation. Humans form complex norms and social groups of remarkable scale compared to other animals. Even within the human experience, the accumulation of knowledge over time has been associated with reduced rates of violence (Pinker 2007) and increases in the scope of cooperation (Wright 2001), from band to tribe to city-state to nation to transnational organization. One might generalize from this trend and argue that as machines approach and exceed human cognitive capacities, Fox, Joshua, and Carl Shulman. 2010. “Superintelligence Does Not Imply Benevolence.” In ECAP10: VIII European Conference on Computing and Philosophy, edited by Klaus Mainzer. Munich: Verlag Dr. Hut. This version contains minor changes. moral behavior will improve in tandem. We argue that this picture neglects a critical distinction between two conceptions of morality, and a related distinction between routes from increased intelligence to more moral behavior. One conception frames morality as a system for cooperation between entities with diverse aims and the ability to affect one another’s pursuit of those aims. Practices such as reciprocal altruism (Trivers 1971), help partners increase their respective reproductive fitnesses. In the cooperative conception, the reason to perform moral behaviors, or to dispose oneself to do so (Gauthier 1986), is to advance one’s own ends. Another, axiological, conception holds that morality demands revision of our ultimate ends. This conception is especially important for treatment of the helpless, e.g., nonhuman animals. Cooperative moral theories, e.g., Gauthier (1986), often can only derive moral status for the helpless from cooperation with altruistic powerful agents. We can then evaluate alternative paths from intelligence to moral behavior. First, machines with greater instrumental rationality could better devise and implement cooperative practices. Thus Hall (2007) argues that intelligent machines will out-cooperate humans, at least with powerful peers. Second, on a Kantian view of morality, one might think that as intelligent machines expanded their knowledge and capacities, they would be directly motivated to revise their preferences to be more moral (Chalmers 2010). We consider a particular counterexample to the Kantian view. Using a definition of intelligence as ability to achieve goals in a wide range of environments (Legg 2008), we discuss the AIXI formalism, which combines Solomonoff induction with Bayesian decision theory to optimize for unknown reward functions (Hutter 2005). AIXI, although physically unrealizable, is a compactly specified superintelligence, provably optimal in maximizing towards arbitrary goals, but has “no room” for the Kantian revision. Instead, it would preserve arbitrary values in most situations (Omohundro 2008). Thus we have reason to think that diverse intelligent machines would convergently display a “drive” to cooperation with sufficiently powerful partners for instrumental reasons, even if this was not specifically engineered. Yet we have reason for pessimism about the ultimate ends of intelligent machines not carefully engineered to be altruistic, and so should work to avoid situations in which such systems are very powerful relative to humanity (Yudkowsky 2008). Joshua Fox, Carl Shulman
DOI: 10.1037/0033-295x.108.4.814
发表时间: 2001-10
影响因子: 5.4
作者:
J. Haidt
通讯作者: J. Haidt