From Bayesian Sparsity to Gated Recurrent Nets

From Bayesian Sparsity to Gated Recurrent Nets
复制标题

DOI:
--
复制
发表时间:
2017-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Hao He;Bo Xin;Satoshi Ikehata;D. Wipf
Hao He;Bo Xin;Satoshi Ikehata;D. Wipf
中科院分区:
其他
文献类型:
--
作者:
Hao He;Bo Xin;Satoshi Ikehata;D. Wipf

文献摘要

被引文献

相似文献

当用于最小化常见正则化回归函数时,许多一阶算法的迭代通常类似于具有预先指定权重的神经网络层。这一观察结果促进了基于学习的方法的发展,这些方法旨在用从可用训练数据中伪造的DNN模型的增强替代这些迭代。例如,重要的NP-hard稀疏估计问题最近从这种类型的升级中受益,简单的前馈或循环网络淘汰了基于近端梯度的迭代。类似地,本文证明了用于提高稀疏性的更强大的贝叶斯算法,它依赖于复杂的多环最大化最小化技术,反映了更复杂的长短期记忆(LSTM)网络的结构,或先前设计用于序列预测的备选门控反馈网络。作为这一发展的一部分,我们研究了在优化过程中跨多个时间尺度运行的潜在变量轨迹之间的相似之处,以及设计用于自适应建模这些特征序列的深度网络结构中的激活。由此产生的见解导致了一种新的稀疏估计系统,当获得训练数据时,该系统可以在其他算法失败的情况下有效地估计最优解,包括实际到达方向(DOA)和3D几何恢复问题。我们揭示的基本原理也暗示了其他领域中更丰富的多循环算法的学习过程。
The iterations of many first-order algorithms, when applied to minimizing common regularized regression functions, often resemble neural network layers with pre-specified weights. This observation has prompted the development of learning-based approaches that purport to replace these iterations with enhanced surrogates forged as DNN models from available training data. For example, important NP-hard sparse estimation problems have recently benefitted from this genre of upgrade, with simple feedforward or recurrent networks ousting proximal gradient-based iterations. Analogously, this paper demonstrates that more powerful Bayesian algorithms for promoting sparsity, which rely on complex multi-loop majorization-minimization techniques, mirror the structure of more sophisticated long short-term memory (LSTM) networks, or alternative gated feedback networks previously designed for sequence prediction. As part of this development, we examine the parallels between latent variable trajectories operating across multiple time-scales during optimization, and the activations within deep network structures designed to adaptively model such characteristic sequences. The resulting insights lead to a novel sparse estimation system that, when granted training data, can estimate optimal solutions efficiently in regimes where other algorithms fail, including practical direction-of-arrival (DOA) and 3D geometry recovery problems. The underlying principles we expose are also suggestive of a learning process for a richer class of multi-loop algorithms in other domains.