Data-driven diagnosis for compressed sensing with cross validation

Data-driven diagnosis for compressed sensing with cross validation
复制标题

DOI:
10.1103/physreve.98.052120
复制
发表时间:
2018-11
期刊:
影响因子:
2.4
通讯作者:
Y. Nakanishi-Ohno;K. Hukushima
Y. Nakanishi-Ohno;K. Hukushima
中科院分区:
物理与天体物理3区
文献类型:
--
作者:
Y. Nakanishi-Ohno;K. Hukushima

文献摘要

相似文献

我们提出了一个数据驱动的过程,用于诊断压缩感知的结果是成功还是失败。压缩感知是实验物理中广泛应用的一种有效的数据采集方法。许多以前的研究表明,压缩感知工作良好,如果稀疏建模可以应用,也就是说,感兴趣的物理现象被假定为仅由几个解释性因素来描述。然而,这是很难确认的假设稀疏建模的各个实例,因为我们不知道的真实表示或其稀疏性的实例正在调查。为了克服这个困难,我们研究了一种称为交叉验证的统计工具,其中所有可用的数据被随机分为两个子集,并且通过仅从另一个子集中的数据估计的稀疏表示来描述一个子集中的数据的准确性被评估为错误。特别是,我们专注于交叉验证误差对两个子集的大小比的依赖性。受统计力学的启发,我们的分析表明,当可用数据的总量正好处于成功与失败之间的临界点时,交叉验证误差渐近地遵循大小比的幂律。因此,可以通过将交叉验证误差的行为与幂律进行比较来以数据驱动的方式诊断压缩感测结果。
We propose a data-driven procedure for diagnosing the results of compressed sensing as success or failure. Compressed sensing is an efficient data-acquisition method widely used in experimental physics. Many previous studies demonstrated that compressed sensing works well if sparse modeling can be applied; that is, physical phenomena of interest are assumed to be described by only a few explanatory factors. However, it is difficult to confirm the assumption of sparse modeling for respective instances in advance because we do not know the true representation or its sparseness of the instance being investigated. To overcome this difficulty, we examined a statistical tool called cross validation, in which all available data are randomly divided into two subsets and the accuracy of describing data in one subset by a sparse representation estimated from only data in the other subset is evaluated as an error. In particular, we focused on the dependence of the cross-validation error on the size ratio of the two subsets. Our analysis, inspired by statistical mechanics, showed that the cross-validation error asymptotically follows a power law of the size ratio when the total amount of available data is exactly at the critical point between success and failure. Hence, compressed-sensing results can be diagnosed in a data-driven manner by comparing the behavior of the cross-validation error with the power law.