Exploring Computational User Models for Agent Policy Summarization

Exploring Computational User Models for Agent Policy Summarization
复制标题

DOI:
10.24963/ijcai.2019/194
复制
发表时间:
2019-05
期刊:
IJCAI : proceedings of the conference
影响因子:
--
通讯作者:
Isaac Lage;Daphna Lifschitz;F. Doshi-Velez;Ofra Amir
Isaac Lage;Daphna Lifschitz;F. Doshi-Velez;Ofra Amir
中科院分区:
其他
文献类型:
--
作者:
Isaac Lage;Daphna Lifschitz;F. Doshi-Velez;Ofra Amir

文献摘要

被引文献

相似文献

人工智能代理支持从驾驶汽车到开药的高风险决策过程,这使得人类用户了解他们的行为变得越来越重要。策略总结方法的目的是通过展示这些代理在信息状态子集中的行为来传达它们的优点和缺点。一些策略总结方法提取总结,该总结在假设用户将部署反向强化学习的情况下优化重构代理策略的能力。在本文中,我们将探讨使用不同的模型来提取摘要。我们引入了一种基于模仿学习的策略摘要方法,通过计算模拟,我们证明了用于提取摘要的模型和用于重建策略的模型之间的不匹配会导致重建质量更差;我们通过一项以人为对象的研究表明,人们在不同的环境中使用不同的模型来重建政策,并且将概要提取模型与这些匹配可以提高性能。总之,我们的研究结果表明,在策略总结中仔细考虑用户模型是很重要的。
AI agents support high stakes decision-making processes from driving cars to prescribing drugs, making it increasingly important for human users to understand their behavior. Policy summarization methods aim to convey strengths and weaknesses of such agents by demonstrating their behavior in a subset of informative states. Some policy summarization methods extract a summary that optimizes the ability to reconstruct the agent's policy under the assumption that users will deploy inverse reinforcement learning. In this paper, we explore the use of different models for extracting summaries. We introduce an imitation learning-based approach to policy summarization; we demonstrate through computational simulations that a mismatch between the model used to extract a summary and the model used to reconstruct the policy results in worse reconstruction quality; and we demonstrate through a human-subject study that people use different models to reconstruct policies in different contexts, and that matching the summary extraction model to these can improve performance. Together, our results suggest that it is important to carefully consider user models in policy summarization.