Mixture of experts: a literature survey

Mixture of experts: a literature survey
复制标题

DOI:
10.1007/s10462-012-9338-y
复制
发表时间:
2014-08-01
影响因子:
12
通讯作者:
Ebrahimpour, Reza
Ebrahimpour, Reza
中科院分区:
计算机科学2区
文献类型:
--
作者:
Masoudnia, Saeed;Ebrahimpour, Reza

文献摘要

被引文献

相似文献

专家混合(ME)是最流行和最有趣的组合方法之一,在提高机器学习性能方面具有巨大的潜力。ME是建立在分而治之的原则,其中的问题空间划分之间的一些神经网络专家,监督的门控网络。在ME的早期作品中,开发了不同的策略来划分专家之间的问题空间。为了更清楚地调查和分析这些方法,我们根据这种差异对ME文献进行了分类。根据所使用的划分策略以及门控网络如何以及何时参与划分和组合过程,将各种ME实现分为两组。在第一组中,传统的ME和这种方法的扩展随机分区的问题空间成若干子空间,使用一个特殊的就业误差函数,专家成为专门在每个子空间。在第二组中,在专家训练过程开始之前,通过聚类方法显式地划分问题空间,然后将每个专家分配到这些子空间中的一个。基于使用专家之间隐性竞争过程的隐式问题空间划分,我们将第一组称为隐式本地化专家混合物(MILE),第二组称为显式本地化专家混合物(MELE),因为它使用了预先指定的集群。两组的性质进行了研究,相互比较。对MILE和MELE的调查,讨论了每一组的优点和缺点,表明这两种方法具有互补的特点。此外,ME方法的特点进行了比较与其他流行的组合方法,包括升压和负相关学习方法。由于所研究的方法具有互补的优势和局限性,以前的研究,试图结合联合收割机的特点,在集成的方法进行了审查,并提出了一些建议,为未来的研究方向。
Mixture of experts (ME) is one of the most popular and interesting combining methods, which has great potential to improve performance in machine learning. ME is established based on the divide-and-conquer principle in which the problem space is divided between a few neural network experts, supervised by a gating network. In earlier works on ME, different strategies were developed to divide the problem space between the experts. To survey and analyse these methods more clearly, we present a categorisation of the ME literature based on this difference. Various ME implementations were classified into two groups, according to the partitioning strategies used and both how and when the gating network is involved in the partitioning and combining procedures. In the first group, The conventional ME and the extensions of this method stochastically partition the problem space into a number of subspaces using a special employed error function, and experts become specialised in each subspace. In the second group, the problem space is explicitly partitioned by the clustering method before the experts' training process starts, and each expert is then assigned to one of these sub-spaces. Based on the implicit problem space partitioning using a tacit competitive process between the experts, we call the first group the mixture of implicitly localised experts (MILE), and the second group is called mixture of explicitly localised experts (MELE), as it uses pre-specified clusters. The properties of both groups are investigated in comparison with each other. Investigation of MILE versus MELE, discussing the advantages and disadvantages of each group, showed that the two approaches have complementary features. Moreover, the features of the ME method are compared with other popular combining methods, including boosting and negative correlation learning methods. As the investigated methods have complementary strengths and limitations, previous researches that attempted to combine their features in integrated approaches are reviewed and, moreover, some suggestions are proposed for future research directions.