Identifiability-Guaranteed Simplex-Structured Post-Nonlinear Mixture Learning via Autoencoder

Identifiability-Guaranteed Simplex-Structured Post-Nonlinear Mixture Learning via Autoencoder
复制标题

DOI:
10.1109/tsp.2021.3096806
复制
发表时间:
2021-06
影响因子:
5.4
通讯作者:
Qi Lyu;Xiao Fu
Qi Lyu;Xiao Fu
中科院分区:
工程技术1区
文献类型:
--
作者:
Qi Lyu;Xiao Fu

文献摘要

相似文献

这项工作的重点是在无监督的方式解开非线性混合的潜在成分的问题。潜在成分被假定为驻留在概率单纯形,并通过一个未知的后非线性混合系统进行转换。该问题在信号和数据分析中有各种应用,例如,非线性高光谱分解、图像嵌入和非线性聚类。线性混合学习问题已经是不适定的,因为目标潜在成分的可识别性通常很难建立。由于涉及未知的非线性,该问题更具挑战性。先前的工作提供了一个函数方程为基础的配方可证明的潜在成分识别。然而,可识别性条件有些严格和不切实际。此外,可识别性分析是基于无限样本(即,总体)案例,而对实际有限样本案例的理解一直难以捉摸。此外,在现有的工作中的算法交易模型的表达性与计算的方便性,这往往阻碍了学习性能。我们的贡献是三方面的。首先,在很大程度上放松的假设下,新的可识别性条件。第二,全面的样本复杂性的结果,这是第一种。第三,提出了一种基于约束自动编码器的算法框架,有效地规避了现有算法的挑战。合成和真实的实验证实了我们的理论分析。
This work focuses on the problem of unraveling nonlinearly mixed latent components in an unsupervised manner. The latent components are assumed to reside in the probability simplex, and are transformed by an unknown post-nonlinear mixing system. This problem finds various applications in signal and data analytics, e.g., nonlinear hyperspectral unmixing, image embedding, and nonlinear clustering. Linear mixture learning problems are already ill-posed, as identifiability of the target latent components is hard to establish in general. With unknown nonlinearity involved, the problem is even more challenging. Prior work offered a function equation-based formulation for provable latent component identification. However, the identifiability conditions are somewhat stringent and unrealistic. In addition, the identifiability analysis is based on the infinite sample (i.e., population) case, while the understanding for practical finite sample cases has been elusive. Moreover, the algorithm in the prior work trades model expressiveness with computational convenience, which often hinders the learning performance. Our contribution is threefold. First, new identifiability conditions are derived under largely relaxed assumptions. Second, comprehensive sample complexity results are presented—which are the first of the kind. Third, a constrained autoencoder-based algorithmic framework is proposed for implementation, which effectively circumvents the challenges in the existing algorithm. Synthetic and real experiments corroborate our theoretical analyses.