Local Similarity Imputation Based on Fast Clustering for Incomplete Data in Cyber-Physical Systems

Local Similarity Imputation Based on Fast Clustering for Incomplete Data in Cyber-Physical Systems
复制标题

DOI:
10.1109/jsyst.2016.2576026
复制
发表时间:
2018-06
影响因子:
4.4
通讯作者:
Liang Zhao;Zhikui Chen;Zhennan Yang;Yueming Hu;M. Obaidat
Liang Zhao;Zhikui Chen;Zhennan Yang;Yueming Hu;M. Obaidat
中科院分区:
计算机科学2区
文献类型:
--
作者:
Liang Zhao;Zhikui Chen;Zhennan Yang;Yueming Hu;M. Obaidat

文献摘要

被引文献

相似文献

由于各种原因,例如传感器故障、通信故障、环境干扰和人为错误,丢失值在信息物理系统(CPS)中很常见。准确的缺失值填补对于提高数据挖掘和统计分析任务的数据质量至关重要。然而,现有的方法大多是利用整个数据集来估算缺失值,这可能会对不相关记录的估算结果产生不利的影响(准确性低或复杂性高)。针对这一问题,本文提出了一种新的局部相似性填补方法,估计缺失数据的快速聚类和前$k$最近邻的基础上。为了提高插补的准确性,一个两层堆叠的自动编码器结合独特的插补应用于定位数据集的主要特征进行聚类。然后,采用前k-最近邻混合距离加权填补算法来填补聚类中的缺失值。通过与现有的四种高质量插补方法进行比较,对五种流行的加州尔湾大学数据集和一种从CPS收集的空气质量监测数据集进行了评估。实验结果表明,该方法可以有效地填补缺失数据值,特别是在CPS的局部特征的不完整数据。
Missing values are common in cyber-physical systems (CPS) for a variety of reasons, such as sensor faults, communication malfunctions, environmental interferences, and human errors. An accurate missing value imputation is crucial to promote the data quality for data mining and statistical analysis tasks. Unfortunately, most of the existing methods take use of the whole data set to impute a missing value, which could have unfavorable influences and impacts (low accuracy or high complexity) on the imputed results caused by irrelevant records. Aiming at this problem, this paper develops a novel local similarity imputation method that estimates missing data based on fast clustering and top $k$-nearest neighbors. To improve the imputation accuracy, a two-layer stacked autoencoder combined with distinctive imputation is applied to locate the principal features of a dataset for clustering. Then, the top $k$ -nearest neighbor hybrid distance weighted imputation is approached to fill in missing values in clusters. The proposed method is evaluated on five popular University of California Irvine datasets as well as one air quality monitoring dataset collected from CPS through comparison with four high-quality existing imputation methods. Empirical results present that the proposed scheme can impute the missing data values effectively and efficiently, especially for the incomplete data with local characteristic in CPS.