Understanding The Robustness of Self-supervised Learning Through Topic Modeling

Understanding The Robustness of Self-supervised Learning Through Topic Modeling
复制标题

DOI:
--
复制
发表时间:
2022-02
期刊:
--
影响因子:
--
通讯作者:
Zeping Luo;Shiyou Wu;C. Weng;Mo Zhou;Rong Ge
Zeping Luo;Shiyou Wu;C. Weng;Mo Zhou;Rong Ge
中科院分区:
其他
文献类型:
--
作者:
Zeping Luo;Shiyou Wu;C. Weng;Mo Zhou;Rong Ge

文献摘要

相似文献

自监督学习显着提高了许多 NLP 任务的性能。然而,自监督学习如何发现有用的表示,以及为什么它比概率模型等传统方法更好,仍然很大程度上未知。在本文中,我们重点关注主题建模的背景,并强调自监督学习的一个关键优势——当应用于主题模型生成的数据时,自监督学习可以忽略特定模型,因此不易受到模型错误指定的影响。特别是,我们证明了基于重建或对比样本的常用自监督目标都可以为一般主题模型恢复有用的后验信息。根据经验,我们表明,相同的目标可以与使用正确模型的后验推理相媲美,同时优于使用错误指定模型的后验推理。
Self-supervised learning has significantly improved the performance of many NLP tasks. However, how can self-supervised learning discover useful representations, and why is it better than traditional approaches such as probabilistic models are still largely unknown. In this paper, we focus on the context of topic modeling and highlight a key advantage of self-supervised learning - when applied to data generated by topic models, self-supervised learning can be oblivious to the specific model, and hence is less susceptible to model misspecification. In particular, we prove that commonly used self-supervised objectives based on reconstruction or contrastive samples can both recover useful posterior information for general topic models. Empirically, we show that the same objectives can perform on par with posterior inference using the correct model, while outperforming posterior inference using misspecified models.