Optimal prediction in the linearly transformed spiked model

Optimal prediction in the linearly transformed spiked model
复制标题

DOI:
10.1214/19-aos1819
复制
发表时间:
2017-09
期刊:
The Annals of Statistics
影响因子:
--
通讯作者:
Edgar Dobriban-;W. Leeb;A. Singer
Edgar Dobriban-;W. Leeb;A. Singer
中科院分区:
其他
文献类型:
--
作者:
Edgar Dobriban-;W. Leeb;A. Singer

文献摘要

被引文献

相似文献

我们考虑线性变换尖峰模型,其中观测值$Y_i$是感兴趣的未观测信号的噪声线性变换$X_i$:\Begin{Align*}Y_i=A_i X_i+\varepsilon_i,\end{Align*}当$i=1,\ldots,n$时。还观察到了变换矩阵$A_i$。我们将$X_i$建模为位于未知低维空间的随机向量。我们应该如何预测未观测信号(回归系数)$X_I$?由于噪声大,对每个观测值分别进行回归的幼稚方法是不准确的。相反,我们发展了最优的线性经验贝叶斯方法,通过跨不同样本的“借款强度”来预测美元XI$。我们的方法适用于大数据集,并且依赖于弱矩假设。分析基于随机矩阵理论。我们讨论了在信号处理、反卷积、低温电子显微镜和高噪声区域丢失数据方面的应用。对于丢失的数据,我们的仿真结果表明,我们的方法比著名的矩阵补全方法更快,对噪声和不等采样更稳健。
We consider the linearly transformed spiked model, where observations $Y_i$ are noisy linear transforms of unobserved signals of interest $X_i$: \begin{align*} Y_i = A_i X_i + \varepsilon_i, \end{align*} for $i=1,\ldots,n$. The transform matrices $A_i$ are also observed. We model $X_i$ as random vectors lying on an unknown low-dimensional space. How should we predict the unobserved signals (regression coefficients) $X_i$? The naive approach of performing regression for each observation separately is inaccurate due to the large noise. Instead, we develop optimal linear empirical Bayes methods for predicting $X_i$ by "borrowing strength" across the different samples. Our methods are applicable to large datasets and rely on weak moment assumptions. The analysis is based on random matrix theory. We discuss applications to signal processing, deconvolution, cryo-electron microscopy, and missing data in the high-noise regime. For missing data, we show in simulations that our methods are faster, more robust to noise and to unequal sampling than well-known matrix completion methods.