A too-good-to-be-true prior to reduce shortcut reliance.

A too-good-to-be-true prior to reduce shortcut reliance.
复制标题

DOI:
10.1016/j.patrec.2022.12.010
复制
发表时间:
2023-02
影响因子:
5.1
通讯作者:
Love, Bradley C.
Love, Bradley C.
中科院分区:
计算机科学3区
文献类型:
--
作者:
Dagaev, Nikolay;Roads, Brett D.;Luo, Xiaoliang;Barry, Daniel N.;Patil, Kaustubh R.;Love, Bradley C.

文献摘要

参考文献

被引文献

相似文献

具有挑战性的机器学习问题不太可能有微不足道的解决方案。来自低容量车型的解决方案很可能是不能推广的捷径。强健概括的一个归纳偏差是避免过于简单的解决方案。低产能模式可以找出捷径,帮助培养高产能模式。尽管深度网络在标准测试条件下的目标识别和其他任务中的表现令人印象深刻,但它们往往不能推广到分布外(o.o.d.)样本。这一缺陷的一个原因是,现代体系结构倾向于依赖与类别相关的ǣ快捷方式ǥ表面功能,而没有捕获跨上下文保持的更深层次的不变量。现实世界的概念通常具有复杂的结构,可以在不同的上下文中表面上有所不同,这可以使一个上下文中最直观和最有希望的解决方案不会推广到其他上下文中。提高o.o.d的一个潜在方法。泛化是假设简单的解决方案不太可能在不同的上下文中有效,并避免它们,我们称之为太好而不真实的先验。具有浅层架构的低容量网络(LCN)应该只能学习表面关系,包括捷径。我们发现,LCN可以作为捷径探测器。此外,LCN的预测可以在两阶段方法中使用,以鼓励高容量网络(HCN)依赖更深层次的不变特征,这些特征应该得到广泛推广。具体地说,当训练HCN时,LCN可以掌握的项目被降低权重。使用我们引入快捷方式的CIFAR-10数据集的修改版本,我们发现两阶段LCN-HCN方法减少了对快捷方式的依赖,并促进了o.o.d。泛化。
Challenging machine learning problems are unlikely to have trivial solutions. Solutions from low-capacity models are likely shortcuts that won’t generalize. One inductive bias for robust generalization is to avoid overly simple solutions. A low-capacity model can identify shortcuts to help train a high-capacity model. Despite their impressive performance in object recognition and other tasks under standard testing conditions, deep networks often fail to generalize to out-of-distribution (o.o.d.) samples. One cause for this shortcoming is that modern architectures tend to rely on ǣshortcutsǥ superficial features that correlate with categories without capturing deeper invariants that hold across contexts. Real-world concepts often possess a complex structure that can vary superficially across contexts, which can make the most intuitive and promising solutions in one context not generalize to others. One potential way to improve o.o.d. generalization is to assume simple solutions are unlikely to be valid across contexts and avoid them, which we refer to as the too-good-to-be-true prior. A low-capacity network (LCN) with a shallow architecture should only be able to learn surface relationships, including shortcuts. We find that LCNs can serve as shortcut detectors. Furthermore, an LCN’s predictions can be used in a two-stage approach to encourage a high-capacity network (HCN) to rely on deeper invariant features that should generalize broadly. In particular, items that the LCN can master are downweighted when training the HCN. Using a modified version of the CIFAR-10 dataset in which we introduced shortcuts, we found that the two-stage LCN-HCN approach reduced reliance on shortcuts and facilitated o.o.d. generalization.
DOI: 10.1016/j.visres.2020.04.013
发表时间: 2020-09-01
期刊: VISION RESEARCH
影响因子: 1.8
作者:
Malhotra, Gaurav;Evans, Benjamin D.;Bowers, Jeffrey S.
通讯作者: Bowers, Jeffrey S.