Overcoming Blind Spots in the Real World: Leveraging Complementary Abilities for Joint Execution

Overcoming Blind Spots in the Real World: Leveraging Complementary Abilities for Joint Execution
复制标题

DOI:
10.1609/aaai.v33i01.33016137
复制
发表时间:
2019-07
期刊:
--
影响因子:
--
通讯作者:
Ramya Ramakrishnan;Ece Kamar;Besmira Nushi;Debadeepta Dey;J. Shah;E. Horvitz
Ramya Ramakrishnan;Ece Kamar;Besmira Nushi;Debadeepta Dey;J. Shah;E. Horvitz
中科院分区:
其他
文献类型:
--
作者:
Ramya Ramakrishnan;Ece Kamar;Besmira Nushi;Debadeepta Dey;J. Shah;E. Horvitz

文献摘要

被引文献

相似文献

模拟器越来越多地用于在将代理部署到真实世界环境中之前对其进行培训。虽然模拟训练提供了一种具有成本效益的学习方式,但模拟器的建模方面不佳可能导致代价高昂的错误或盲点。虽然人类可以帮助引导智能体识别这些错误区域,但人类本身在执行中也存在盲点和噪音。我们研究了如何学习的盲点都可以用来管理交接的决定时,人类和代理人共同行动,在现实世界中,他们都没有得到充分的训练或评估。该公式假设代理盲点来自模拟世界中的代表性限制,这导致代理忽略与开放世界中的行为相关的重要特征。我们的盲点发现方法将模拟中收集的经验与有限的人类演示相结合。第一步,将模仿学习应用于演示数据,以识别人类正在使用但智能体缺失的重要特征。第二步使用从仿真和演示数据中的智能体和人类之间的动作不匹配中提取的噪声标签来训练盲点模型。我们通过两个领域的实验表明,我们的方法是能够学习一个简洁的表示,准确地捕捉盲点区域,并避免危险的错误,在真实的世界中,通过代理和人类之间的控制转移。
Simulators are being increasingly used to train agents before deploying them in real-world environments. While training in simulation provides a cost-effective way to learn, poorly modeled aspects of the simulator can lead to costly mistakes, or blind spots. While humans can help guide an agent towards identifying these error regions, humans themselves have blind spots and noise in execution. We study how learning about blind spots of both can be used to manage hand-off decisions when humans and agents jointly act in the real-world in which neither of them are trained or evaluated fully. The formulation assumes that agent blind spots result from representational limitations in the simulation world, which leads the agent to ignore important features that are relevant for acting in the open world. Our approach for blind spot discovery combines experiences collected in simulation with limited human demonstrations. The first step applies imitation learning to demonstration data to identify important features that the human is using but that the agent is missing. The second step uses noisy labels extracted from action mismatches between the agent and the human across simulation and demonstration data to train blind spot models. We show through experiments on two domains that our approach is able to learn a succinct representation that accurately captures blind spot regions and avoids dangerous errors in the real world through transfer of control between the agent and the human.