How Well Do Unsupervised Learning Algorithms Model Human Real-time and Life-long Learning?

How Well Do Unsupervised Learning Algorithms Model Human Real-time and Life-long Learning?
复制标题

DOI:
--
复制
发表时间:
2022
期刊:
Advances in neural information processing systems
影响因子:
--
通讯作者:
Chengxu Zhuang;Ziyu Xiang;Yoon Bai;Xiaoxuan Jia;N. Turk-Browne;K. Norman;J. DiCarlo;Daniel Yamins
Chengxu Zhuang;Ziyu Xiang;Yoon Bai;Xiaoxuan Jia;N. Turk-Browne;K. Norman;J. DiCarlo;Daniel Yamins
中科院分区:
其他
文献类型:
--
作者:
Chengxu Zhuang;Ziyu Xiang;Yoon Bai;Xiaoxuan Jia;N. Turk-Browne;K. Norman;J. DiCarlo;Daniel Yamins

文献摘要

相似文献

人类在多个时间尺度上从视觉输入中学习,既可以在短时间内快速灵活地获取视觉知识,又可以在较长时间内稳健地积累在线学习进展。对这些强大的学习能力进行建模是计算视觉认知科学的一个重要问题,而能够复制它们的模型将在现实世界的计算机视觉环境中发挥重要作用。在这项工作中,我们建立了实时和终身持续视觉学习的基准。我们的实时学习基准测试衡量了模型在给定视觉输入流的情况下,在数分钟和数小时的过程中匹配真实的人类的快速视觉行为变化的能力。我们的终身学习基准评估模型在纯在线学习课程中的表现,这些课程直接从儿童多年发展过程中的视觉经验中获得。我们在这两个基准上评估了一系列最近的深度自监督视觉学习算法,发现它们都没有完全匹配人类的表现,尽管有些算法的表现比其他算法好得多。有趣的是,体现自我监督学习最新趋势的算法-包括BYOL,SwAV和MAE -在我们的基准测试中比上一代的自我监督算法(如Simplified和MoCo-v2)要差得多。我们目前的分析表明,这些新的算法的失败主要是由于他们无法处理的那种稀疏的低多样性的数据流,自然出现在真实的世界,并积极利用内存通过负采样-这些新的算法回避的机制-似乎有利于在这种低多样性的环境中学习。我们还说明了两个基准测试中的短时间尺度和长时间尺度之间的互补性,展示了如何要求单个学习算法对本地上下文足够敏感以匹配实时学习变化,同时足够稳定以避免长期的灾难性遗忘,这会导致类似人类的算法可能必须跨越的权衡。总之,我们的基准建立了一种定量的方法来直接比较神经网络模型和人类学习者之间的学习,显示了这些算法处理样本比较和记忆的机制中的选择如何强烈影响它们与人类学习能力相匹配的能力,并为识别更灵活和强大的视觉自我监督算法提供了一个开放的问题空间。
Humans learn from visual inputs at multiple timescales, both rapidly and flexibly acquiring visual knowledge over short periods, and robustly accumulating online learning progress over longer periods. Modeling these powerful learning capabilities is an important problem for computational visual cognitive science, and models that could replicate them would be of substantial utility in real-world computer vision settings. In this work, we establish benchmarks for both real-time and life-long continual visual learning. Our real-time learning benchmark measures a model's ability to match the rapid visual behavior changes of real humans over the course of minutes and hours, given a stream of visual inputs. Our life-long learning benchmark evaluates the performance of models in a purely online learning curriculum obtained directly from child visual experience over the course of years of development. We evaluate a spectrum of recent deep self-supervised visual learning algorithms on both benchmarks, finding that none of them perfectly match human performance, though some algorithms perform substantially better than others. Interestingly, algorithms embodying recent trends in self-supervised learning - including BYOL, SwAV and MAE - are substantially worse on our benchmarks than an earlier generation of self-supervised algorithms such as SimCLR and MoCo-v2. We present analysis indicating that the failure of these newer algorithms is primarily due to their inability to handle the kind of sparse low-diversity datastreams that naturally arise in the real world, and that actively leveraging memory through negative sampling - a mechanism eschewed by these newer algorithms - appears useful for facilitating learning in such low-diversity environments. We also illustrate a complementarity between the short and long timescales in the two benchmarks, showing how requiring a single learning algorithm to be locally context-sensitive enough to match real-time learning changes while stable enough to avoid catastrophic forgetting over the long term induces a trade-off that human-like algorithms may have to straddle. Taken together, our benchmarks establish a quantitative way to directly compare learning between neural networks models and human learners, show how choices in the mechanism by which such algorithms handle sample comparison and memory strongly impact their ability to match human learning abilities, and expose an open problem space for identifying more flexible and robust visual self-supervision algorithms.