Incorporating Behavioral Constraints in Online AI Systems

Incorporating Behavioral Constraints in Online AI Systems
复制标题

将行为约束纳入在线人工智能系统

DOI:
10.1609/aaai.v33i01.33013
复制
发表时间:
2018
期刊:
影响因子:
0.6
通讯作者:
F. Rossi
F. Rossi
中科院分区:
物理与天体物理4区
文献类型:
--
作者:
Avinash Balakrishnan;Djallel Bouneffouf;Nicholas Mattei;F. Rossi

文献摘要

被引文献

相似文献

通过奖励反馈来学习其所采取行动的人工智能系统越来越多地部署在对我们日常生活有重大影响的领域。然而,在许多情况下,在线奖励不应成为唯一的指导标准,因为法规、价值观、偏好或道德原则还施加了额外的限制和/或优先事项。我们详细介绍了一种新颖的在线代理,它通过观察学习一组行为约束,并在在线环境中做出决策时使用这些学习到的约束作为指导,同时仍然对奖励反馈做出反应。为了定义这个代理,我们建议对经典的上下文多臂老虎机设置采用一种新颖的扩展,并提供一种称为行为约束汤普森采样(BCTS)的新算法,该算法允许在遵守外生约束的同时进行在线学习。我们的代理学习一个约束策略,该策略实现教师代理所展示的观察到的行为约束,然后使用该约束策略来指导基于奖励的在线探索和利用。我们描述了我们代理背后的上下文强盗算法的预期遗憾的上限,并提供了两个应用领域中真实世界数据的案例研究。我们的实验表明,设计的智能体能够在一组行为约束内行动,而不会显着降低其整体奖励绩效。
AI systems that learn through reward feedback about the actions they take are increasingly deployed in domains that have significant impact on our daily life. However, in many cases the online rewards should not be the only guiding criteria, as there are additional constraints and/or priorities imposed by regulations, values, preferences, or ethical principles. We detail a novel online agent that learns a set of behavioral constraints by observation and uses these learned constraints as a guide when making decisions in an online setting while still being reactive to reward feedback. To define this agent, we propose to adopt a novel extension to the classical contextual multi-armed bandit setting and we provide a new algorithm called Behavior Constrained Thompson Sampling (BCTS) that allows for online learning while obeying exogenous constraints. Our agent learns a constrained policy that implements the observed behavioral constraints demonstrated by a teacher agent, and then uses this constrained policy to guide the reward-based online exploration and exploitation. We characterize the upper bound on the expected regret of the contextual bandit algorithm that underlies our agent and provide a case study with real world data in two application domains. Our experiments show that the designed agent is able to act within the set of behavior constraints without significantly degrading its overall reward performance.