From “Thumbs Up” to “10 out of 10”: Reconsidering Scalar Feedback in Interactive Reinforcement Learning

From “Thumbs Up” to “10 out of 10”: Reconsidering Scalar Feedback in Interactive Reinforcement Learning
复制标题

DOI:
10.1109/iros55552.2023.10342458
复制
发表时间:
2023-10
期刊:
2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
影响因子:
--
通讯作者:
Hang Yu;Reuben M. Aronson;Katherine H. Allen;E. Short
Hang Yu;Reuben M. Aronson;Katherine H. Allen;E. Short
中科院分区:
其他
文献类型:
--
作者:
Hang Yu;Reuben M. Aronson;Katherine H. Allen;E. Short

文献摘要

相似文献

从人类反馈中学习是提高机器人在探索任务中学习能力的有效途径。与二进制人类反馈的广泛应用相比,标量人类反馈被认为是噪声和不稳定的,因此使用较少。在本文中,我们比较了标量和二进制反馈,并证明了标量反馈有利于学习时,妥善处理。我们收集了二进制或标量反馈,分别从两组机器人任务的众工。我们发现,当考虑参与者如何一致地标记相同的数据时,标量反馈导致的一致性低于二进制反馈;然而,如果允许小的不匹配,差异就会消失。此外,标量和二进制反馈在与关键强化学习目标的相关性方面没有显着差异。然后,我们引入了稳定教师评估动态(STEADY),以改善从标量反馈的学习。基于标量反馈是多分布的思想,STEADY重构了潜在的正反馈和负反馈分布,并基于反馈统计对标量反馈进行了重新标度。我们表明,在机器人达到非专家人类反馈的任务中,使用标量反馈+ STEADY训练的模型优于基线,包括二进制反馈和原始标量反馈。我们的研究结果表明,二进制反馈和标量反馈都是动态的,标量反馈是一个很有前途的信号,用于交互式强化学习。
Learning from human feedback is an effective way to improve robotic learning in exploration-heavy tasks. Compared to the wide application of binary human feedback, scalar human feedback has been used less because it is believed to be noisy and unstable. In this paper, we compare scalar and binary feedback, and demonstrate that scalar feedback benefits learning when properly handled. We collected binary or scalar feedback respectively from two groups of crowdworkers on a robot task. We found that when considering how consistently a participant labeled the same data, scalar feedback led to less consistency than binary feedback; however, the difference vanishes if small mismatches are allowed. Additionally, scalar and binary feedback show no significant differences in their correlations with key Reinforcement Learning targets. We then introduce Stabilizing TEacher Assessment DYnamics (STEADY) to improve learning from scalar feedback. Based on the idea that scalar feedback is muti-distributional, STEADY reconstructs underlying positive and negative feedback distributions and re-scales scalar feedback based on feedback statistics. We show that models trained with scalar feedback + STEADY outperform baselines, including binary feedback and raw scalar feedback, in a robot reaching task with non-expert human feedback. Our results show that both binary feedback and scalar feedback are dynamic, and scalar feedback is a promising signal for use in interactive Reinforcement Learning.