Meta-Learning Operators to Optimality from Multi-Task Non-IID Data

Meta-Learning Operators to Optimality from Multi-Task Non-IID Data
复制标题

DOI:
10.48550/arxiv.2308.04428
复制
发表时间:
2023-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Thomas Zhang;Leonardo F. Toso;James Anderson;N. Matni
Thomas Zhang;Leonardo F. Toso;James Anderson;N. Matni
中科院分区:
其他
文献类型:
--
作者:
Thomas Zhang;Leonardo F. Toso;James Anderson;N. Matni

文献摘要

相似文献

机器学习最新进展的一个强大概念是从异构来源或任务中提取跨数据的共同特征。直观地,使用一个人的所有数据来学习通用表示功能,从而使计算工作和统计概括都受益于较小的参数以微调给定的任务。为了理论上扎根这些优点,我们提出了从嘈杂的矢量测量值$ y = mx + w $中恢复线性运算符$ m $的一般环境,其中covariates $ x $可能是non-i.i.i.d。和非偏性。我们证明,现有的各向同性 - 敏捷的元学习方法在表示更新中产生偏见,这会导致噪声项的缩放缩放,从而失去对源任务数量的良好依赖性。反过来,这可能会导致表示学习的样本复杂性被单任务数据大小瓶颈。我们引入了一种改编,$ \ texttt {de-bias&feature-whiten} $($ \ texttt {dfw} $),这是Collins等人(2021)中提出的流行交替交替的最小化 - 淡淡(AMD)方案(2021),并与噪声级别的最佳代表建立了lineare contrangence in lineare contional side suild suilds $ suild condits $。这导致与Oracle经验风险最小化器相同的顺序构成概括。我们验证了各种数值模拟上$ \ texttt {dfw} $的至关重要性。特别是,我们表明香草交替的最小化下降即使对于IID但轻度的非异端数据也会灾难性地失败。我们的分析统一并概括了先前的工作,并为更广泛的应用程序(例如在控件和动态系统中)提供了灵活的框架。
A powerful concept behind much of the recent progress in machine learning is the extraction of common features across data from heterogeneous sources or tasks. Intuitively, using all of one's data to learn a common representation function benefits both computational effort and statistical generalization by leaving a smaller number of parameters to fine-tune on a given task. Toward theoretically grounding these merits, we propose a general setting of recovering linear operators $M$ from noisy vector measurements $y = Mx + w$, where the covariates $x$ may be both non-i.i.d. and non-isotropic. We demonstrate that existing isotropy-agnostic meta-learning approaches incur biases on the representation update, which causes the scaling of the noise terms to lose favorable dependence on the number of source tasks. This in turn can cause the sample complexity of representation learning to be bottlenecked by the single-task data size. We introduce an adaptation, $\texttt{De-bias&Feature-Whiten}$ ($\texttt{DFW}$), of the popular alternating minimization-descent (AMD) scheme proposed in Collins et al., (2021), and establish linear convergence to the optimal representation with noise level scaling down with the $\textit{total}$ source data size. This leads to generalization bounds on the same order as an oracle empirical risk minimizer. We verify the vital importance of $\texttt{DFW}$ on various numerical simulations. In particular, we show that vanilla alternating-minimization descent fails catastrophically even for iid, but mildly non-isotropic data. Our analysis unifies and generalizes prior work, and provides a flexible framework for a wider range of applications, such as in controls and dynamical systems.