Real vs. simulated: Questions on the capability of simulated datasets on building fault detection for energy efficiency from a data-driven perspective

Real vs. simulated: Questions on the capability of simulated datasets on building fault detection for energy efficiency from a data-driven perspective
复制标题

DOI:
10.1016/j.enbuild.2022.111872
复制
发表时间:
2022-02-02
影响因子:
6.7
通讯作者:
Candan, Kasim Selcuk
Candan, Kasim Selcuk
中科院分区:
工程技术2区
文献类型:
--
作者:
Huang, Jiajing;Wen, Jin;Candan, Kasim Selcuk

文献摘要

被引文献

相似文献

由于建筑物真实的数据获取和分析的难度大、费用高,建筑物故障自动检测与诊断(AFDD)的研究主要集中在模拟系统数据上。缺乏对使用模拟数据的数据驱动AFDD方法的性能和可扩展性的验证,以及如何将其与来自真实的建筑数据的性能和可扩展性进行比较。在这项研究中,我们进行了两组实验来寻求这个问题的答案。我们首先评估数据驱动的故障检测策略的真实的和模拟的建筑数据分别。我们观察到,故障检测性能不受故障检测策略,训练数据的大小,以及交叉验证次数的影响时,训练和盲测试数据来自相同的数据源,即,模拟或真实的建筑数据。接下来,我们进行了跨数据集的研究,也就是说,使用模拟数据开发模型,并在真实的建筑数据上进行测试。结果表明,在模拟数据上训练的模型不能推广应用于真实的建筑物数据的故障检测。进行Kolmogorov-Smirnov检验,以确认模拟和真实的建筑数据之间存在统计差异,并识别两个数据集之间具有相似性的特征子集。使用该特征的子集,跨数据集实验表明,故障检测在大多数故障情况下的改进。我们的结论是,即使系统从物理分析的角度产生具有相同故障症状的模拟数据,并非模拟数据集的所有特征都可能对AFDD不利,但从机器学习的角度来看,只有一部分特征包含有价值的信息。(c)2022 Elsevier B. V.保留所有权利。
Literature on building Automatic Fault Detection and Diagnosis (AFDD) mainly focuses on simulated system data due to high expenses and difficulties of obtaining and analyzing real building data. There is a lack of validation on performances and scalabilities of data-driven AFDD approaches using simulated data and how it compares to that from real building data. In this study, we conduct two sets of experiments to seek answers to this question. We first evaluate data-driven fault detection strategies on real and simulated building data separately. We observe that the fault detection performances are not affected by fault detection strategies, sizes of training data, and the number of cross-validation folds when training and blind test data come from the same data source, namely, simulated or real building data. Next, we conduct a cross-dataset study, that is, develop the model using simulated data and tested on real building data. The results indicate the model trained on simulated data is not generalized to be applied for real building data for fault detection. Kolmogorov-Smirnov Test is conducted to confirm that there exist statistical differences between the simulated and real building data and identify a subset of features with similarities between the two datasets. Using the subset of the feature, cross-dataset experiments show fault detection improvements on most fault cases. We conclude that even if the system produces simulated data with the same fault symptoms from physical analysis perspectives, not all features from simulated datasets may not be beneficial for AFDD but only a subset of features contains valuable information from a machine learning perspective. (c) 2022 Elsevier B.V. All rights reserved.