Dealing with missing data in MSPC: several methods, different interpretations, some examples

Dealing with missing data in MSPC: several methods, different interpretations, some examples
复制标题

DOI:
10.1002/cem.750
复制
发表时间:
2002-08-01
影响因子:
2.4
通讯作者:
Ferrer, A
Ferrer, A
中科院分区:
化学3区
文献类型:
--
作者:
Arteaga, F;Ferrer, A

文献摘要

被引文献

相似文献

本文解决了使用未来的缺失数据的多变量观察来估计现有主成分分析 (PCA) 模型的潜在变量分数的问题。这是多元统计过程控制 (MSPC) 方案中的一个关键问题,其中基于基础 PCA 模型不断询问过程。我们提出了几种估计缺失数据的新个体得分的方法:所谓的修剪得分法(TRI)、单分量投影法(SCP)、投影到模型平面的方法(PMP)、基于缺失数据迭代插补的方法、基于平方预测误差最小化的方法(SPE)、条件均值替换法(CMR)和各种基于最小二乘的方法:基于已知数据回归的方法(KDR)另一个基于修剪分数(TSR)的回归。开发了每种方法的基础以及分数估计器的表达式、它们的协方差矩阵和估计误差。所讨论的一些方法已经在文献中提出(SCP、PMP 和 CMR),一些是原创的(TRI 和 TSR),另一些则被证明与其他作者已经开发的方法等效:迭代插补和 SPE 方法与 PMP 等效; KDR相当于CMR。这些方法可以被视为估算缺失变量值的不同方法。通过基于工业数据集的模拟研究了这些方法的效率。 KDR 方法在统计上优于其他方法,但 TSR 方法除外,其中要求逆的矩阵的尺寸要小得多。版权所有 (C) 2002 John Wiley Sons, Ltd.
This paper addresses the problem of using future multivariate observations with missing data to estimate latent variable scores from an existing principal component analysis (PCA) model. This is a critical issue in multivariate statistical process control (MSPC) schemes where the process is continuously interrogated based on an underlying PCA model. We present several methods for estimating the scores of new individuals with missing data: a so-called trimmed score method (TRI), a single-component projection method (SCP), a method of projection to the model plane (PMP), a method based on the iterative imputation of missing data, a method based on the minimization of the squared prediction error (SPE), a conditional mean replacement method (CMR) and various least squared-based methods: one based on a regression on known data (KDR) and the other based on a regression on trimmed scores (TSR). The basis for each method and the expressions for the score estimators, their covariance matrices and the estimation errors are developed. Some of the methods discussed have already been proposed in the literature (SCP, PMP and CMR), some are original (TRI and TSR) and others are shown to be equivalent to methods already developed by other authors: iterative imputation and SPE methods are equivalent to PMP; KDR is equivalent to CMR. These methods can be seen as different ways to impute values for the missing variables. The efficiency of the methods is studied through simulations based on an industrial data set. The KDR method is shown to be statistically superior to the other methods, except the TSR method in which the matrix to be inverted is of a much smaller size. Copyright (C) 2002 John Wiley Sons, Ltd.