Making Human-Like Trade-offs in Constrained Environments by Learning from Demonstrations

Making Human-Like Trade-offs in Constrained Environments by Learning from Demonstrations
复制标题

通过从演示中学习,在受限环境中做出类似人类的权衡

DOI:
--
复制
发表时间:
2021
期刊:
arXiv.org
影响因子:
--
通讯作者:
K. Venable
K. Venable
中科院分区:
--
文献类型:
--
作者:
Arie Glazier;Andrea Loreggia;Nicholas Mattei;Taher Rahgooy;F. Rossi;K. Venable

文献摘要

参考文献

被引文献

相似文献

许多现实生活中的场景要求人类做出艰难的权衡:我们是总是遵守所有的交通规则,还是在紧急情况下违反限速?这些情景迫使我们评估集体规范和我们自己的个人目标之间的权衡。为了创建有效的人工智能团队,我们必须为人工智能代理提供一个模型,说明人类如何在复杂、受限的环境中进行权衡。这些智能体将能够反映人类的行为,或者将人类的注意力吸引到可以改善决策的情况。为此,我们提出了一种新的逆强化学习(IRL)方法,用于从演示中学习隐式硬约束和软约束,使代理能够快速适应新的设置。此外,学习状态、动作和状态特征的软约束允许代理将这些知识转移到共享相似方面的新领域。然后,我们使用约束学习方法来实现一个新的系统架构,利用人类决策的认知模型,多选择决策场理论(MDFT),编排竞争的目标。我们评估所产生的代理的轨迹长度,违反限制的数量,和总奖励,证明我们的代理架构是通用的,并实现了强大的性能。因此,我们能够捕捉和复制类似人类的权衡,从环境中的约束条件不明确的演示。
Many real-life scenarios require humans to make difficult trade-offs: do we always follow all the traffic rules or do we violate the speed limit in an emergency? These scenarios force us to evaluate the trade-off between collective norms and our own personal objectives. To create effective AI-human teams, we must equip AI agents with a model of how humans make trade-offs in complex, constrained environments. These agents will be able to mirror human behavior or to draw human attention to situations where decision making could be improved. To this end, we propose a novel inverse reinforcement learning (IRL) method for learning implicit hard and soft constraints from demonstrations, enabling agents to quickly adapt to new settings. In addition, learning soft constraints over states, actions, and state features allows agents to transfer this knowledge to new domains that share similar aspects. We then use the constraint learning method to implement a novel system architecture that leverages a cognitive model of human decision making, multi-alternative decision field theory (MDFT), to orchestrate competing objectives. We evaluate the resulting agent on trajectory length, number of violated constraints, and total reward, demonstrating that our agent architecture is both general and achieves strong performance. Thus we are able to capture and replicate human-like trade-offs from demonstrations in environments when constraints are not explicit.
DOI: --
发表时间: 2018
期刊: Thirty-third Conference on Neural Information Processing Systems (NeurIPS
影响因子: --
作者:
VazquezChanlatte, Marcell;Jha, Susmit;Tiwari, Ashish;Seshia, Sanjit
通讯作者: Seshia, Sanjit
DOI: 10.1037//0033-295x.108.2.370
发表时间: 2001-04-01
影响因子: 5.4
作者:
Roe, RM;Busemeyer, JR;Townsend, JT
通讯作者: Townsend, JT