A Parsimonious Mixture of Gaussian Trees Model for Oversampling in Imbalanced and Multimodal Time-Series Classification

A Parsimonious Mixture of Gaussian Trees Model for Oversampling in Imbalanced and Multimodal Time-Series Classification
复制标题

DOI:
10.1109/tnnls.2014.2308321
复制
发表时间:
2014-12-01
影响因子:
10.4
通讯作者:
Pang, John Z. F.
Pang, John Z. F.
中科院分区:
计算机科学1区
文献类型:
--
作者:
Cao, Hong;Tan, Vincent Y. F.;Pang, John Z. F.

文献摘要

被引文献

相似文献

我们提出了一个新的框架,使用一个简约的统计模型,被称为混合高斯树,建模可能的多模态少数类,以解决不平衡的时间序列分类的问题。通过利用由于时间序列的平滑性而导致的邻近时间点高度相关的事实,我们的模型将待估计的协方差参数的数量从O(d(2))显著减少到O(Ld),其中L是混合成分的数量,d是维度。因此,我们的模型是特别有效的建模高维时间序列的少数正类的实例数量有限。此外,学习模型的计算复杂度仅为O(Ln+d(2))阶,其中n(+)是正标记样本的数量。我们基于几个著名的时间序列数据集(单模态和多模态)进行了广泛的分类实验,首先从我们学习的混合模型中随机生成合成实例,以纠正不平衡。然后,我们将我们的结果与几种最先进的过采样技术进行比较,结果表明,当我们提出的模型用于过采样时,相同的支持向量机分类器在整个数据集范围内实现了更好的分类精度。事实上,所提出的方法实现了最好的平均性能30倍的36个多模态数据集根据F值度量。我们的研究结果也具有很强的竞争力相比,nonoversampling为基础的分类器处理不平衡的时间序列数据集。
We propose a novel framework of using a parsimonious statistical model, known as mixture of Gaussian trees, for modeling the possibly multimodal minority class to solve the problem of imbalanced time-series classification. By exploiting the fact that close-by time points are highly correlated due to smoothness of the time-series, our model significantly reduces the number of covariance parameters to be estimated from O(d(2)) to O(Ld), where L is the number of mixture components and d is the dimensionality. Thus, our model is particularly effective for modeling high-dimensional time-series with limited number of instances in the minority positive class. In addition, the computational complexity for learning the model is only of the order O(Ln+d(2)) where n(+) is the number of positively labeled samples. We conduct extensive classification experiments based on several well-known time-series data sets (both single- and multimodal) by first randomly generating synthetic instances from our learned mixture model to correct the imbalance. We then compare our results with several state-of-the-art oversampling techniques and the results demonstrate that when our proposed model is used in oversampling, the same support vector machines classifier achieves much better classification accuracy across the range of data sets. In fact, the proposed method achieves the best average performance 30 times out of 36 multimodal data sets according to the F-value metric. Our results are also highly competitive compared with nonoversampling-based classifiers for dealing with imbalanced time-series data sets.