Generating Deontic Obligations From Utility-Maximizing Systems

Generating Deontic Obligations From Utility-Maximizing Systems
复制标题

从效用最大化系统生成道义义务

DOI:
10.1145/3514094.3534163
复制
发表时间:
2022
期刊:
Ethics and Society
影响因子:
--
通讯作者:
Abbas, Houssam
Abbas, Houssam
中科院分区:
--
文献类型:
--
作者:
Shea-Blymyer, Colin;Abbas, Houssam

文献摘要

参考文献

相似文献

这项工作给出了一个经过强化学习(RL)训练的代理的(伦理和社会)义务的逻辑特征。RL代理通过遵循效用最大化策略来采取行动。我们认为,效用函数的选择隐含着伦理和社会价值,有必要明确这些价值。这项工作为这样做提供了基础。首先,我们提出了一种概率道义逻辑,它适合于形式化地规定随机系统的义务,包括它的伦理义务。我们证明了这一逻辑的一些有用的有效性,以及它的语义如何与马尔可夫决策过程(MDP)的语义兼容。其次,我们证明了模型检查允许我们证明一个代理有一个给定的义务来实现某种事件状态--这意味着通过最优的行动,它正在寻求达到这种事件状态。我们开发了一个模型检查器,用于针对MDP的逻辑。第三,我们观察到,对于系统设计者来说,获得其系统义务的逻辑特征是有用的,这可能比效用函数的表达更易于解释,并且在调试中更有帮助。列举代理的所有义务是不切实际的,因此我们提出了一个贝叶斯优化例程,该例程学习生成系统设计者认为感兴趣的系统义务。我们实现了模型检验和贝叶斯优化例程,并通过初步的试点研究证明了它们的有效性。这项工作提供了一种严格的方法,根据他们隐含地寻求满足的(伦理和社会)义务来表征效用最大化的代理人。
This work gives a logical characterization of the (ethical and social) obligations of an agent trained with Reinforcement Learning (RL). An RL agent takes actions by following a utility-maximizing policy. We maintain that the choice of utility function embeds ethical and social values implicitly, and that it is necessary to make these values explicit. This work provides a basis for doing so. First, we propose a probabilistic deontic logic that is suited for formally specifying the obligations of a stochastic system, including its ethical obligations. We prove some useful validities about this logic, and how its semantics are compatible with those of Markov Decision Processes (MDPs). Second, we show that model checking allows us to prove that an agent has a given obligation to bring about some state of affairs - meaning that by acting optimally, it is seeking to reach that state of affairs. We develop a model checker for our logic against MDPs. Third, we observe that it is useful for a system designer to obtain a logical characterization of her system's obligations, which is potentially more interpretable and helpful in debugging than the expression of a utility function. Enumerating all the obligations of an agent is impractical, so we propose a Bayesian optimization routine that learns to generate a system's obligations that the system designer deems interesting. We implement the model checking and Bayesian optimization routines, and demonstrate their effectiveness with an initial pilot study. This work provides a rigorous method to characterize utility-maximizing agents in terms of the (ethical and social) obligations that they implicitly seek to satisfy.
算法伦理:自动驾驶汽车义务的形式化和验证
DOI: 10.1145/3460975
发表时间: 2021
期刊: ACM Transactions on Cyber-Physical Systems (TCPS)
影响因子: --
作者:
Colin Shea;Houssam Abbas
通讯作者: Houssam Abbas
规范多智能体系统的计算模型
DOI: --
发表时间: 2013
期刊: Normative Multi-Agent Systems
影响因子: --
作者:
N. Alechina;Nick Bassiliades;M. Dastani;Marina De Vos;B. Logan;S. Mera;Andreasa Morris;F. Schapachnik
通讯作者: F. Schapachnik
祈使句的逻辑形式
DOI: --
发表时间: 1975
期刊:
影响因子: --
作者:
D. Clarke
通讯作者: D. Clarke
通过机械化道义逻辑迈向道德机器人*
DOI: --
发表时间: 2005
期刊:
影响因子: --
作者:
Konstantine Arkoudas;S. Bringsjord;Paul Bello
通讯作者: Paul Bello