Using the EM algorithm to train neural networks: misconceptions and a new algorithm for multiclass classification

Using the EM algorithm to train neural networks: misconceptions and a new algorithm for multiclass classification
复制标题

DOI:
10.1109/tnn.2004.826217
复制
发表时间:
2004-05
影响因子:
--
通讯作者:
S. Ng;G. McLachlan
S. Ng;G. McLachlan
中科院分区:
--
文献类型:
--
作者:
S. Ng;G. McLachlan

文献摘要

被引文献

相似文献

期望最大化(EM)算法作为神经网络应用领域(如模式识别)中各种算法的基础,近年来引起了相当大的兴趣。然而,关于它在神经网络中的应用存在一些误解。在本文中,我们澄清这些误解,并考虑如何EM算法可以通过训练多层感知器(MLP)和混合专家(ME)网络的应用程序,多类分类。我们确定了一些情况下,EM算法训练MLP网络的应用可能是有限的价值,并讨论了一些处理困难的方法。对于ME网络,文献中报道,在M步的内环中使用迭代重加权最小二乘(IRLS)算法的EM算法训练的网络在多类分类中通常表现不佳。然而,我们发现,IRLS算法的收敛是稳定的,对数似然是单调增加的学习率小于1时,采用。此外,我们建议使用的期望条件最大化(ECM)算法来训练ME网络。在一些模拟和真实的数据集上,它的性能被证明是上级的IRLS算法。
The expectation-maximization (EM) algorithm has been of considerable interest in recent years as the basis for various algorithms in application areas of neural networks such as pattern recognition. However, there exists some misconceptions concerning its application to neural networks. In this paper, we clarify these misconceptions and consider how the EM algorithm can be adopted to train multilayer perceptron (MLP) and mixture of experts (ME) networks in applications to multiclass classification. We identify some situations where the application of the EM algorithm to train MLP networks may be of limited value and discuss some ways of handling the difficulties. For ME networks, it is reported in the literature that networks trained by the EM algorithm using iteratively reweighted least squares (IRLS) algorithm in the inner loop of the M-step, often performed poorly in multiclass classification. However, we found that the convergence of the IRLS algorithm is stable and that the log likelihood is monotonic increasing when a learning rate smaller than one is adopted. Also, we propose the use of an expectation-conditional maximization (ECM) algorithm to train ME networks. Its performance is demonstrated to be superior to the IRLS algorithm on some simulated and real data sets.