Large scale multi-output multi-class classification using Gaussian processes

Large scale multi-output multi-class classification using Gaussian processes
复制标题

DOI:
10.1007/s10994-022-06289-3
复制
发表时间:
2023-02
期刊:
影响因子:
7.5
通讯作者:
Chunchao Ma;Mauricio A Álvarez
Chunchao Ma;Mauricio A Álvarez
中科院分区:
计算机科学3区
文献类型:
--
作者:
Chunchao Ma;Mauricio A Álvarez

文献摘要

相似文献

多输出高斯过程(MOGP)通过利用与其他输出变量的相关性,可以帮助提高对某些输出变量的预测性能。在本文中,我们的主要动机是使用多输出高斯过程来挖掘输出之间的相关性,其中每一输出都是一个多类分类问题。MOGP多用于多输出回归。有一些现有的工作将MOGP用于其他类型的输出,例如,多输出二进制分类。然而,用于多类分类的MOGP研究较少。原因有两个:1)当使用Softmax函数时,不清楚如何将其扩展到几个输出的情况之外;2)多类分类问题中最常见的数据类型包括图像数据,并且MOGP不是专门为图像数据设计的。因此,我们提出了一种新的多输出高斯过程模型,称为增减多输出高斯过程(MOGPS-AR),该模型能够处理大规模分类和降维图像输入数据。大规模分类是通过对每个输出中的训练数据集和类进行二次采样来实现的,而缩小尺寸的图像输入数据是通过在新模型中加入卷积核来处理的。我们的实验表明,无论是在合成问题上还是在实际分类问题上,我们提出的模型在不同的性能度量方面都优于单输出高斯过程,在可伸缩性方面优于多输出高斯过程。我们使用OmMiclot数据集提供了一个示例,在其中我们展示了我们模型的属性。
Multi-output Gaussian processes (MOGPs) can help to improve predictive performance for some output variables, by leveraging the correlation with other output variables. In this paper, our main motivation is to use multiple-output Gaussian processes to exploit correlations between outputs where each output is a multi-class classification problem. MOGPs have been mostly used for multi-output regression. There are some existing works that use MOGPs for other types of outputs, e.g., multi-output binary classification. However, MOGPs for multi-class classification has been less studied. The reason is twofold: 1) when using a softmax function, it is not clear how to scale it beyond the case of a few outputs; 2) most common type of data in multi-class classification problems consists of image data, and MOGPs are not specifically designed to image data. We thus propose a new MOGPs model calledMulti-output Gaussian Processes with Augment & Reduce (MOGPs-AR)that can deal with large scale classification and downsized image input data. Large scale classification is achieved by subsampling both training data sets and classes in each output whereas downsized image input data is handled by incorporating a convolutional kernel into the new model. We show empirically that our proposed model outperforms single-output Gaussian processes in terms of different performance metrics and multi-output Gaussian processes in terms of scalability, both in synthetic and in real classification problems. We include an example with the Ommiglot dataset where we showcase the properties of our model.