Learning Models of Sequential Decision-Making with Partial Specification of Agent Behavior

Learning Models of Sequential Decision-Making with Partial Specification of Agent Behavior
复制标题

具有部分代理行为规范的顺序决策学习模型

DOI:
--
复制
发表时间:
2019
期刊:
AAAI Conference on Artificial Intelligence
影响因子:
--
通讯作者:
J. Shah
J. Shah
中科院分区:
--
文献类型:
--
作者:
Vaibhav Unhelkar;J. Shah

文献摘要

被引文献

相似文献

与其他(人类或人工)代理交互的人工代理需要模型来推理其他代理的行为。除了这些模型的预测效用之外,保持与代理的真实行为生成模型一致的模型对于有效的人-代理交互至关重要。在应用程序中,观察和部分规范的代理的行为是可用的,实现模型对齐是具有挑战性的各种原因。首先,代理人的决策因素往往是不完全知道的;此外,以前的方法,只依赖于代理人的行为的观察可能无法恢复真正的模型,因为多个模型可以解释观察到的行为同样好。为了实现更好的模型对齐,我们提供了一种新的方法,能够学习对齐的模型,符合代理的行为的部分知识。我们的方法的核心是一个因子化的行为模型(AMM),沿着贝叶斯非参数先验,以及能够将部分规范作为模型学习的约束的推理方法。我们在实验中评估我们的方法,并证明模型对齐的指标的改进。
Artificial agents that interact with other (human or artificial) agents require models in order to reason about those other agents’ behavior. In addition to the predictive utility of these models, maintaining a model that is aligned with an agent’s true generative model of behavior is critical for effective human-agent interaction. In applications wherein observations and partial specification of the agent’s behavior are available, achieving model alignment is challenging for a variety of reasons. For one, the agent’s decision factors are often not completely known; further, prior approaches that rely upon observations of agents’ behavior alone can fail to recover the true model, since multiple models can explain observed behavior equally well. To achieve better model alignment, we provide a novel approach capable of learning aligned models that conform to partial knowledge of the agent’s behavior. Central to our approach are a factored model of behavior (AMM), along with Bayesian nonparametric priors, and an inference approach capable of incorporating partial specifications as constraints for model learning. We evaluate our approach in experiments and demonstrate improvements in metrics of model alignment.