Analysis of Large Heterogeneous Repairable System Reliability Data with Static System Attributes and Dynamic Sensor Measurement in Big Data Environment

Analysis of Large Heterogeneous Repairable System Reliability Data with Static System Attributes and Dynamic Sensor Measurement in Big Data Environment
复制标题

大数据环境下具有静态系统属性和动态传感器测量的大型异构可修复系统可靠性数据分析

DOI:
10.1080/00401706.2019.1609584
复制
发表时间:
2019
期刊:
影响因子:
2.5
通讯作者:
Pan, Rong
Pan, Rong
中科院分区:
工程技术3区
文献类型:
--
作者:
Liu, Xiao;Pan, Rong

文献摘要

相似文献

在大数据时代,工程师面临的一个紧迫挑战是对大量具有协变量的异构可修复系统进行可靠性分析。除了静态协变量(包括时不变系统属性,如标称操作条件、地理位置等)之外,传感技术的最新进展也使得获得系统操作和环境条件的动态传感器测量成为可能。在大数据环境中,大量的可靠性数据通常存储在分布式存储系统中。利用现代统计学习的力量,本文研究了一种统计方法,它集成了随机森林算法和经典的数据分析方法,如平均累积函数的非参数估计和基于非齐次泊松过程的参数模型的可修系统可靠性。我们表明,所提出的方法有效地解决了一些常见的挑战,从实践中产生的,包括系统异构性,协变量的选择,模型规范和数据的局部性,由于分布式数据存储。建立了估计量的大样本性质和一致相合性。两个数值例子和案例研究来说明所提出的方法的应用。所提出的方法的优点是通过比较研究证明。数据集和计算机代码已在GitHub上提供。
In the age of Big Data, one pressing challenge facing engineers is to perform reliability analysis for a large fleet of heterogeneous repairable systems with covariates. In addition to static covariates, which include time-invariant system attributes such as nominal operating conditions, geo-locations, etc., the recent advances of sensing technologies have also made it possible to obtain dynamic sensor measurement of system operating and environmental conditions. As a common practice in the Big Data environment, the massive reliability data are typically stored in some distributed storage systems. Leveraging the power of modern statistical learning, this article investigates a statistical approach which integrates the random forests algorithm and the classical data analysis methodologies for repairable system reliability, such as the nonparametric estimator for the mean cumulative function and the parametric models based on the nonhomogeneous Poisson process. We show that the proposed approach effectively addresses some common challenges arising from practice, including system heterogeneity, covariate selection, model specification and data locality due to the distributed data storage. The large sample properties as well as the uniform consistency of the proposed estimator are established. Two numerical examples and a case study are presented to illustrate the application of the proposed approach. The strengths of the proposed approach are demonstrated by comparison studies. Datasets and computer code have been made available on GitHub.