A Novel Semiparametric Hidden Markov Model for Process Failure Mode Identification

A Novel Semiparametric Hidden Markov Model for Process Failure Mode Identification
复制标题

DOI:
10.1109/tase.2016.2636292
复制
发表时间:
2018-04
影响因子:
5.6
通讯作者:
Hongyang Yu
Hongyang Yu
中科院分区:
计算机科学1区
文献类型:
--
作者:
Hongyang Yu

文献摘要

被引文献

相似文献

隐马尔可夫模型(HMM)的发射分布通常使用过程变量的交叉矩来构造。与单变量概率分布的均值类似,交叉矩是多变量概率分布的最基本统计量,它不能捕捉过程数据的高阶统计特征。为了克服这一局限性,本文利用交叉矩的高阶等价性作为完全依赖结构来构造HMM的发射分布。过程变量之间的完整的依赖结构是在高斯copula建模。为了保证使用高斯Copula的必要条件得到满足,还提出了一种半参数数据变换。最终的发射分布被构造为copula模型的有限混合。本文提出的隐马尔可夫模型在两个工业过程中进行了性能验证。从业人员注意-高斯混合模型隐马尔可夫模型(GMM-HMM)是动态工业过程模式识别的一种有效且易于实现的工具。GMM-HMM的这些优点主要来自于高斯发射分布的使用,它允许计算封闭形式的导数用于模式参数的期望和最大化(EM)估计。所提出的Copula混合模型HMM(SPCMM-HMM)的主要新奇是双重的;其发射分布也属于指数族,从而能够通过精确的EM过程有效地估计参数;同时,它也能够表征过程变量之间的复杂依赖结构。然而,与GMM-HMM相比,提出的SPCMM-HMM具有更多的模型参数。它在被监控的过程系统的操作高度集成(具有复杂的过程变量交互)并且有丰富的训练数据样本的情况下表现出色。在本文中,过程变量之间的相互作用在这两个案例研究是复杂的,由于实施多个闭环控制回路,大量的训练数据样本可以很容易地从仿真生成。SPCMM-HMM的性能被证明是一贯优于GMM-HMM在这样的设置下,这是相当普遍的现代工业过程。尽管如此,如果所考虑的过程系统仅设计用于简单操作或训练数据样本稀缺,则GMM-HMM仍然是首选方法,因为它不太容易过拟合。
The emitting distributions of a hidden Markov model (HMM) are normally constructed using the cross moments of the process variables. Similar to the mean of a univariate probability distribution, the cross moment is the most fundamental statistic of a multivariate probability distribution, which is not capable of capturing the high-order statistical features of process data. To alleviate this limitation, the high-order equivalence of the cross moment demonstrated in this paper, as the complete dependence structure, is used to construct the emitting distribution for HMM. The complete dependence structure among the process variables is modeled in a Gaussian copula. A semiparametric data transformation is also proposed to ensure the necessary conditions for using a Gaussian copula are met. The final emitting distribution is constructed as a finite mixture of the copula models. The proposed HMM is tested on two industrial studies for performance validation.Note to Practitioners—Gaussian mixture model HMM (GMM-HMM) is an efficient and easy-to-implement tool for mode identification of dynamic industrial processes. These virtues of GMM-HMM mainly come from the use of Gaussian emitting distribution, which allows closed-form derivatives to be computed for the expectation and maximization (EM) estimation of mode parameters. The main novelty of the proposed copula mixture model HMM (SPCMM-HMM) is twofold; its emitting distribution also belongs to the exponential family, thus enabling efficient estimation of parameters through an exact EM procedure; meanwhile, it is also capable of characterizing complex dependence structures among process variables. However, the proposed SPCMM-HMM has more model parameters compared with the GMM-HMM. It excels in situations where the operation of the process system being monitored is highly integrated (with complex process variable interactions) and abundance of training data samples is available. In this paper, the interactions between process variables in both case studies are complex due to the implementation of multiple closed control loops, and a large amount of training data samples can be easily generated from simulation. The performance of the SPCMM-HMM is shown to be consistently better than that of the GMM-HMM under such settings, which are fairly common in modern industrial processes. Nonetheless, if the process system being considered is only designed for a simple operation or there is a scarcity of training data samples, the GMM-HMM is still the preferred method as it is less prone to overfitting.