An EM-based Ensemble Learning Algorithm on Piecewise Surface Regression Problem

An EM-based Ensemble Learning Algorithm on Piecewise Surface Regression Problem
复制标题

基于EM的分段曲面回归问题集成学习算法

DOI:
--
复制
发表时间:
2012
期刊:
影响因子:
--
通讯作者:
Yuan Li
Yuan Li
中科院分区:
--
文献类型:
--
作者:
Juan Luo;A. Brodsky;Yuan Li

文献摘要

被引文献

相似文献

针对市场细分研究、消费者行为模式识别、气象研究中的天气模式等领域的典型应用,提出了一种基于多步期望最大化(EM-Based)的分段曲面回归算法。所涉及的多个步骤是对训练数据集的每个数据点及其最接近的邻居的一小部分进行局部回归,在由局部回归形成的特征向量空间上进行聚类,对每个单独的表面进行回归学习,以及分类以确定每个单独的表面的边界。在回归学习阶段引入了基于EM的迭代过程,以改善学习结果。在这一阶段,集成学习在为每个数据点重新分配聚类索引方面发挥着重要作用。通过对子模型的预测误差、数据点到回归超平面的距离、数据点到每个聚类曲面质心的距离的多数表决来确定聚类索引的重新分配。在最后执行分类,以确定每个单独曲面的边界。将聚类质量有效性技术应用于输入域的表面数目事先未知的情况。基于人工生成数据源和基准数据源的一系列实验将该算法与广泛使用的回归学习包进行了比较,结果表明,该算法在均方根误差方面优于那些包,尤其是在应用了集成学习之后。
A multi-step Expectation-Maximization based (EM-based) algorithm is proposed to solve the piecewise surface regression problem which has typical applications in market segmentation research, identification of consumer behavior patterns, weather patterns in meteorological research, and so on. The multiple steps involved are local regression on each data point of the training data set and a small set of its closest neighbors, clustering on the feature vector space formed from the local regression, regression learning for each individual surface, and classification to determine the boundaries for each individual surface. An EM-based iteration process is introduced in the regression learning phase to improve the learning outcome. In this phase, ensemble learning plays an important role in the reassignment of the cluster index for each data point. The reassignment of cluster index is determined by the majority voting of predictive error of sub-models, the distance between the data point and regressed hyperplane, and the distance between the data point and centroid of each clustered surface. Classification is performed at the end to determine the boundaries for each individual surface. Clustering quality validity techniques are applied to the scenario in which the number of surfaces for the input domain is not known in advance. A set of experiments based on both artificial generated and benchmark data source are conducted to compare the proposed algorithm and widely-used regression learning packages to show that the proposed algorithm outperforms those packages in terms of root mean squared errors, especially after ensemble learning has been applied.