Is Clustering Advantageous in Statistical Ill-Posed Linear Inverse Problems?

Is Clustering Advantageous in Statistical Ill-Posed Linear Inverse Problems?
复制标题

DOI:
10.1109/tit.2020.3014409
复制
发表时间:
2020-11-01
影响因子:
2.5
通讯作者:
Pensky, Marianna
Pensky, Marianna
中科院分区:
计算机科学2区
文献类型:
--
作者:
Rajapakshage, Rasika;Pensky, Marianna

文献摘要

被引文献

相似文献

在许多统计线性逆问题中,需要在一个没有有界逆的算子下从噪声图像中恢复一类相似的对象。这类问题出现在许多应用领域。通常,在这样的问题中,在预处理步骤中执行聚类,然后分别为每个聚类平均解决逆问题。因此,通常只在估计步骤中检查程序的误差。本文的目的是从理论和模拟两个方面检验聚类对一般不适定线性反问题解的精度的影响。特别地,我们假设X-m=Af(M)+Delta是(M)的元素,m=1,…,M,其中函数f(M)可以分成K个类,并且需要恢复一个向量函数f=(f(1),...,f(M))(T)。我们构造了f的估计量,作为惩罚优化问题的解,它对应于先聚类后估计的设置。我们得到了一个关于它的精度的预言式,并证明了该估计是极小极大最优的或几乎极小极大最优的,直到观测次数的对数倍。我们方法的优点之一是,我们不假设簇的数量是预先知道的。随后,我们将上述方法的精度分别与未聚类的估计精度和恢复每个未知函数后的聚类的估计精度进行了比较。我们的结论是,当问题是适度不适定的时候,在预处理步骤进行聚类是有益的。当问题严重不适时,应极其小心地应用它。
In many statistical linear inverse problems, one needs to recover classes of similar objects from their noisy images under an operator that does not have a bounded inverse. Problems of this kind appear in many areas of application. Routinely, in such problems clustering is carried out at a preprocessing step and then the inverse problem is solved for each of the cluster averages separately. As a result, the errors of the procedures are usually examined for the estimation step only. The objective of this paper is to examine, both theoretically and via simulations, the effect of clustering on the accuracy of the solutions of general ill-posed linear inverse problems. In particular, we assume that one observes X-m = Af(m) + delta is an element of(m), m = 1, ..., M, where functions f(m) can be grouped into K classes and one needs to recover a vector function f = (f(1), ..., f(M))(T). We construct an estimator for f as a solution of a penalized optimization problem which corresponds to the clustering before estimation setting. We derive an oracle inequality for its precision and confirm that the estimator is minimax optimal or nearly minimax optimal up to a logarithmic factor of the number of observations. One of the advantages of our approach is that we do not assume that the number of clusters is known in advance. Subsequently, we compare the accuracy of the above procedure with the precision of estimation without clustering, and clustering following the recovery of each of the unknown functions separately. We conclude that clustering at the pre-processing step is beneficial when the problem is moderately ill-posed. It should be applied with extreme care when the problem is severely ill-posed.