Eigen-Entropy: A metric for multivariate sampling decisions

Eigen-Entropy: A metric for multivariate sampling decisions
复制标题

特征熵:多元采样决策的度量

DOI:
10.1016/j.ins.2022.11.023
复制
发表时间:
2023
影响因子:
8.1
通讯作者:
O'Neill, Zheng
O'Neill, Zheng
中科院分区:
计算机科学1区
文献类型:
--
作者:
Huang, Jiajing;Yoon, Hyunsoo;Wu, Teresa;Candan, Kasim Selcuk;Pradhan, Ojas;Wen, Jin;O'Neill, Zheng

文献摘要

相似文献

采样是一种帮助识别代表性数据子集的技术,该数据子集捕获整个数据集的特征。大多数现有的抽样算法需要对多变量数据进行分布假设,这可能事先无法获得。本文提出了一种基于信息熵的多元数据集特征熵度量方法。EE是一种无模型度量,因为它是基于从相关系数矩阵中提取的特征值导出的,而无需对数据分布进行任何假设。我们证明了EE测量数据集的组成,例如其异质性或同质性。因此,EE可用于支持采样决策,例如针对感兴趣的应用考虑哪些样本和多少样本。为了演示EE度量的实用性,考虑两组用例。第一个用例集中在不平衡数据集的分类问题上,EE用于指导少数类中同质样本的渲染。使用10个公共数据集,它表明,两个过采样技术使用建议EE方法优于文献报道的方法在精度,召回率,F-测量,和G-均值。在第二个实验中,建筑故障检测的EE是用来采样异构数据,以支持故障检测。采用真实的建筑系统的历史正态数据集,对14个测试用例进行了EE基线构建,实验结果表明,EE方法在召回率方面优于基准方法。我们得出结论,EE是一个可行的指标,以支持抽样决策。
Sampling is a technique to help identify a representative data subset that captures the characteristics of the whole dataset. Most existing sampling algorithms require distribution assumptions of the multivariate data, which may not be available beforehand. This study proposes a new metric called Eigen-Entropy (EE), which is based on information entropy for the multivariate dataset. EE is a model-free metric because it is derived based on eigenvalues extracted from the correlation coefficient matrix without any assumptions on data distributions. We prove that EE measures the composition of the dataset, such as its heterogeneity or homogeneity. As a result, EE can be used to support sampling decisions, such as which samples and how many samples to consider with respect to the application of interest. To demonstrate the utility of the EE metric, two sets of use cases are considered. The first use case focuses on classification problems with an imbalanced dataset, and EE is used to guide the rendering of homogeneous samples from minority classes. Using 10 public datasets, it is demonstrated that two oversampling techniques using the proposed EE method outperform reported methods from the literature in terms of precision, recall, F-measure, and G-mean. In the second experiment, building fault detection is investigated where EE is used to sample heterogeneous data to support fault detection. Historical normal datasets collected from real building systems are used to construct the baselines by EE for 14 test cases, and experimental results indicate that the EE method outperforms benchmark methods in terms of recall. We conclude that EE is a viable metric to support sampling decisions.