Imputation Method Based on Collaborative Filtering and Clustering for the Missing Data of the Squeeze Casting Process Parameters

Imputation Method Based on Collaborative Filtering and Clustering for the Missing Data of the Squeeze Casting Process Parameters
复制标题

基于协同过滤和聚类的挤压铸造工艺参数缺失数据插补方法

DOI:
10.1007/s40192-021-00248-x
复制
发表时间:
2022-01-27
影响因子:
3.3
通讯作者:
Liu,Guangming
Liu,Guangming
中科院分区:
材料科学3区
文献类型:
--
作者:
Deng,Jianxin;Ye,Zhixing;Liu,Guangming

文献摘要

相似文献

开发一种高效的方法,从过去的数据建立挤压铸造工艺参数是必不可少的。然而,当存在许多缺失值时,基于过去的数据设计挤压铸造工艺参数是困难的。传统的缺失数据方法在应用于高维多变量缺失数据,特别是具有相关性的材料过程数据时,充满了额外的计算挑战。由于材料成分与工艺参数之间的关系与用户与感兴趣信息之间的关系具有相似的特征,本文提出了一种基于聚类的协同过滤(ClubCF)算法的缺失数据填补方法。将缺失数据样本和非缺失数据样本分为两组,对非缺失数据样本采用基于冠层算法的K-means聚类,得到k个子类数据,然后利用基于Pearson相似度用户填充的协同过滤理论,选择子类数据的值填充缺失数据样本.利用缺失的铝合金挤压铸造工艺参数数据对该方法进行了评价,并进行了更多的对比实验,以了解其性能和特点。采用平均绝对误差和标准差两个指标来量化插补性能,并与三种传统方法(均值插值、回归插值和期望最大化算法)进行比较。结果表明,该方法是有效的,并优于传统的方法处理高维相关数据。
The development of a highly efficient methodology for establishing squeeze casting process parameters from past data is essential. However, designing squeeze casting process parameters based on past data is difficult when there are many missing values. Conventional missing data approaches are fraught with additional computational challenges when applied to high-dimensional multivariable missing data, especially material process data with correlation. As the relationship between material composition and process parameters has similar characteristics with that between users and information of interest, this paper proposes a method for missing data imputation based on a clustering-based collaborative filtering (ClubCF) algorithm to address this challenge. Data samples with and without missing values were divided into two groups.K-means clustering based on a canopy algorithm was applied to the data samples without missing values to obtainksubclass data, whose values were then selected to fill data samples with missing values via a collaborative filtering theory based on Pearson similarity user filling. The missing squeeze casting process parameters data of aluminum alloys were used to evaluate the method, and more comparative experiments were carried out to understand their performance and features. Two different indicators, including the mean absolute error and the standard deviation, were utilized to quantify the imputation performance, which was compared with those of three conventional methods (mean interpolation, regression interpolation, and the expectation maximization algorithm). The results indicate that the proposed approach is effective and outperforms conventional methods in processing high-dimensional correlated data.