Masked Prediction: A Parameter Identifiability View

Masked Prediction: A Parameter Identifiability View
复制标题

DOI:
--
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Bingbin Liu;Daniel J. Hsu;Pradeep Ravikumar;Andrej Risteski
Bingbin Liu;Daniel J. Hsu;Pradeep Ravikumar;Andrej Risteski
中科院分区:
其他
文献类型:
--
作者:
Bingbin Liu;Daniel J. Hsu;Pradeep Ravikumar;Andrej Risteski

文献摘要

相似文献

自我监督学习的绝大多数工作都集中在评估一组选定的下游任务恢复的特征。虽然有几个常用的基准数据集,但这种特征学习镜头需要对下游任务进行假设,而这些假设并非数据分布本身所固有的。在本文中,我们提出了另一种视角,即参数可识别性:假设数据来自参数概率模型,我们用合适的参数形式训练自监督学习预测器,并询问最优预测器的参数是否可以用于提取地面实况生成模型的参数。具体来说,我们专注于捕获序列结构的潜变量模型,即具有离散和条件高斯观测的隐马尔可夫模型。我们将屏蔽预测作为自监督学习任务,并研究最佳屏蔽预测器。我们表明,参数可识别性受到任务难度的控制,任务难度由数据模型的选择和要预测的标记数量决定。在技​​术方面,我们发现了与张量秩分解的独特性的密切联系,张量秩分解是通过矩量法研究可识别性的广泛使用的工具。
The vast majority of work in self-supervised learning have focused on assessing recovered features by a chosen set of downstream tasks. While there are several commonly used benchmark datasets, this lens of feature learning requires assumptions on the downstream tasks which are not inherent to the data distribution itself. In this paper, we present an alternative lens, one of parameter identifiability: assuming data comes from a parametric probabilistic model, we train a self-supervised learning predictor with a suitable parametric form, and ask whether the parameters of the optimal predictor can be used to extract the parameters of the ground truth generative model. Specifically, we focus on latent-variable models capturing sequential structures, namely Hidden Markov Models with both discrete and conditionally Gaussian observations. We focus on masked prediction as the self-supervised learning task and study the optimal masked predictor. We show that parameter identifiability is governed by the task difficulty, which is determined by the choice of data model and the amount of tokens to predict. Technique-wise, we uncover close connections with the uniqueness of tensor rank decompositions , a widely used tool in studying identifiability through the lens of the method of moments.