Truncated variational EM for semi-supervised neural simpletrons

Truncated variational EM for semi-supervised neural simpletrons
复制标题

DOI:
10.1109/ijcnn.2017.7966331
复制
发表时间:
2017-02
期刊:
2017 International Joint Conference on Neural Networks (IJCNN)
影响因子:
--
通讯作者:
D. Forster;Jörg Lücke
D. Forster;Jörg Lücke
中科院分区:
其他
文献类型:
--
作者:
D. Forster;Jörg Lücke

文献摘要

被引文献

相似文献

概率生成网络的推理和学习通常非常具有挑战性,并且通常无法扩展到用于深度判别方法的大型网络。为了获得用于半监督学习的高效可训练、大规模且性能良好的生成网络,我们在这里结合了两项最新进展:分层泊松混合物的神经网络重构(神经简单子)和一种新颖的截断变分EM方法(TV-EM)。 TV-EM 为生成网络中的学习提供了理论保证,其在神经简单子中的应用导致了学习方程的特别紧凑但近似最优的修改。如果应用于标准基准,我们凭经验发现:学习在更少的 EM 迭代中收敛,每次 EM 迭代的复杂性降低,并且最终的似然值平均更高。对于标签很少的数据集的分类任务,与没有截断的应用程序相比,学习改进会导致错误率持续降低。本文对 MNIST 数据集进行的实验可以在半监督环境中与标准模型和最先进的模型进行比较。对 NIST SD19 数据集的进一步实验表明,当有大量额外的未标记数据可用时,该方法具有可扩展性。
Inference and learning for probabilistic generative networks is often very challenging and typically prevents scalability to as large networks as used for deep discriminative approaches. To obtain efficiently trainable, large-scale and well performing generative networks for semi-supervised learning, we here combine two recent developments: a neural network reformulation of hierarchical Poisson mixtures (Neural Simpletrons), and a novel truncated variational EM approach (TV-EM). TV-EM provides theoretical guarantees for learning in generative networks, and its application to Neural Simpletrons results in particularly compact, yet approximately optimal, modifications of learning equations. If applied to standard benchmarks, we empirically find: that learning converges in fewer EM iterations, that the complexity per EM iteration is reduced, and that final likelihood values are higher on average. For the task of classification on data sets with few labels, learning improvements result in consistently lower error rates if compared to applications without truncation. Experiments on the MNIST data set herein allow for comparison to standard and state-of-the-art models in the semi-supervised setting. Further experiments on the NIST SD19 data set show the scalability of the approach when a very large amount of additional unlabeled data is available.