Learning Representations that Enable Generalization in Assistive Tasks
Learning Representations that Enable Generalization in Assistive Tasks
复制标题
能够泛化辅助任务的学习表示
DOI:
10.48550/arxiv.2212.03175
复制
发表时间:
2022
期刊:
影响因子:
--
通讯作者:
A. Dragan
中科院分区:
文献类型:
--
作者:
Jerry Zhi;Aditi Raghunathan;Daniel S. Brown;Zackory M. Erickson;A. Dragan
Recent work in sim2real has successfully enabled robots to act in physical environments by training in simulation with a diverse ''population'' of environments (i.e. domain randomization). In this work, we focus on enabling generalization in assistive tasks: tasks in which the robot is acting to assist a user (e.g. helping someone with motor impairments with bathing or with scratching an itch). Such tasks are particularly interesting relative to prior sim2real successes because the environment now contains a human who is also acting. This complicates the problem because the diversity of human users (instead of merely physical environment parameters) is more difficult to capture in a population, thus increasing the likelihood of encountering out-of-distribution (OOD) human policies at test time. We advocate that generalization to such OOD policies benefits from (1) learning a good latent representation for human policies that test-time humans can accurately be mapped to, and (2) making that representation adaptable with test-time interaction data, instead of relying on it to perfectly capture the space of human policies based on the simulated population only. We study how to best learn such a representation by evaluating on purposefully constructed OOD test policies. We find that sim2real methods that encode environment (or population) parameters and work well in tasks that robots do in isolation, do not work well in assistance. In assistance, it seems crucial to train the representation based on the history of interaction directly, because that is what the robot will have access to at test time. Further, training these representations to then predict human actions not only gives them better structure, but also enables them to be fine-tuned at test-time, when the robot observes the partner act. https://adaptive-caregiver.github.io.
DOI:
10.1613/jair.1.13326
发表时间:
2021-07
期刊:
J. Artif. Intell. Res.
影响因子:
--
作者:
Yuexiang Zhai;Christina Baek;Zhengyuan Zhou;Jiantao Jiao;Yi Ma
通讯作者:
Yuexiang Zhai;Christina Baek;Zhengyuan Zhou;Jiantao Jiao;Yi Ma
DOI:
--
发表时间:
2020-11
期刊:
ArXiv
影响因子:
--
作者:
Annie Xie;Dylan P. Losey;R. Tolsma;Chelsea Finn;Dorsa Sadigh
通讯作者:
Annie Xie;Dylan P. Losey;R. Tolsma;Chelsea Finn;Dorsa Sadigh