Learning from non-random data in Hilbert spaces: an optimal recovery perspective

Learning from non-random data in Hilbert spaces: an optimal recovery perspective
复制标题

DOI:
10.1007/s43670-022-00022-w
复制
发表时间:
2022-04
期刊:
Sampling Theory, Signal Processing, and Data Analysis
影响因子:
--
通讯作者:
S. Foucart;Chunyang Liao;Shahin Shahrampour;Yinsong Wang
S. Foucart;Chunyang Liao;Shahin Shahrampour;Yinsong Wang
中科院分区:
其他
文献类型:
--
作者:
S. Foucart;Chunyang Liao;Shahin Shahrampour;Yinsong Wang

文献摘要

相似文献

经典统计学习中的泛化概念通常与数据点是独立且同分布(IID)随机变量的假设相关联。虽然在许多应用中都是相关的,但这一假设可能并不适用于一般情况,因此鼓励开发对非iid数据具有鲁棒性的学习框架。在这项工作中,我们从最优恢复的角度考虑回归问题。依靠一个类似于选择假设类的模型假设,学习者的目标是最小化最坏情况误差,而不依赖于数据的任何概率假设。我们首先开发了一个计算有限维希尔伯特空间中任何恢复映射的最坏情况误差的半确定程序。然后,对于任何希尔伯特空间,我们证明了最优恢复提供了一个公式,从算法的角度来看,它是用户友好的,只要假设类是线性的。有趣的是,在某些情况下,这个公式与核无脊回归相吻合,证明最小化平均误差和最坏情况误差可以产生相同的解。我们提供数值实验来支持我们的理论发现。
The notion of generalization in classical Statistical Learning is often attached to the postulate that data points are independent and identically distributed (IID) random variables. While relevant in many applications, this postulate may not hold in general, encouraging the development of learning frameworks that are robust to non-IID data. In this work, we consider the regression problem from an Optimal Recovery perspective. Relying on a model assumption comparable to choosing a hypothesis class, a learner aims at minimizing the worst-case error, without recourse to any probabilistic assumption on the data. We first develop a semidefinite program for calculating the worst-case error of any recovery map in finite-dimensional Hilbert spaces. Then, for any Hilbert space, we show that Optimal Recovery provides a formula which is user-friendly from an algorithmic point-of-view, as long as the hypothesis class is linear. Interestingly, this formula coincides with kernel ridgeless regression in some cases, proving that minimizing the average error and worst-case error can yield the same solution. We provide numerical experiments in support of our theoretical findings.