Avoiding Lingering in Learning Active Recognition by Adversarial Disturbance

Avoiding Lingering in Learning Active Recognition by Adversarial Disturbance
复制标题

DOI:
10.1109/wacv56688.2023.00459
复制
发表时间:
2023-01
期刊:
2023 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)
影响因子:
--
通讯作者:
Lei Fan;Ying Wu
Lei Fan;Ying Wu
中科院分区:
其他
文献类型:
--
作者:
Lei Fan;Ying Wu

文献摘要

相似文献

本文认为,主动识别的情况下,代理被授权智能地获得更好的识别观察。代理通常由两个模块组成,即,策略和识别器,以选择操作并预测类别。当使用地面实况类标签来监督识别器时,通常会使用由当前训练中识别器确定的奖励来更新策略,例如是否实现正确的预测。然而,这种联合学习过程可能会导致意想不到的解决方案,比如一个崩溃的策略,它只访问识别器已经经过充分训练以获得奖励的视图,这会损害泛化能力。我们称这种现象为挥之不去,以描述代理不愿意在训练过程中探索具有挑战性的观点。现有的方法来解决勘探开发权衡可能是无效的,因为他们通常假设可靠的反馈,在勘探过程中更新的估计很少访问的国家。这种假设在这里是无效的,因为识别器的奖励可能没有得到充分的训练。为此,我们的方法集成了另一个对抗策略,在训练过程中不断干扰识别代理,形成一个竞争游戏,以促进积极的探索,避免逗留。当识别失败时,强化的对手得到奖励,通过将摄像机转向具有挑战性的观察结果来与识别代理进行竞争。在两个数据集上的大量实验验证了所提出的方法在识别性能,学习效率,特别是在管理环境噪声的鲁棒性方面的有效性。
This paper considers the active recognition scenario, where the agent is empowered to intelligently acquire observations for better recognition. The agents usually compose two modules, i.e., the policy and the recognizer, to select actions and predict the category. While using ground-truth class labels to supervise the recognizer, the policy is typically updated with rewards determined by the current in-training recognizer, like whether achieving correct predictions. However, this joint learning process could lead to unintended solutions, like a collapsed policy that only visits views that the recognizer is already sufficiently trained to obtain rewards, which harms the generalization ability. We call this phenomenon lingering to depict the agent being reluctant to explore challenging views during training. Existing approaches to tackle the exploration-exploitation trade-off could be ineffective as they usually assume reliable feedback during exploration to update the estimate of rarely-visited states. This assumption is invalid here as the reward from the recognizer could be insufficiently trained.To this end, our approach integrates another adversarial policy to constantly disturb the recognition agent during training, forming a competing game to promote active explorations and avoid lingering. The reinforced adversary, rewarded when the recognition fails, contests the recognition agent by turning the camera to challenging observations. Extensive experiments across two datasets validate the effectiveness of the proposed approach regarding its recognition performances, learning efficiencies, and especially robustness in managing environmental noises.