When is there a representer theorem? Vector versus matrix regularizers

When is there a representer theorem? Vector versus matrix regularizers
复制标题

DOI:
10.5555/1577069.1755870
复制
发表时间:
2008-09
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
Andreas Argyriou;C. Micchelli;M. Pontil
Andreas Argyriou;C. Micchelli;M. Pontil
中科院分区:
其他
文献类型:
--
作者:
Andreas Argyriou;C. Micchelli;M. Pontil

文献摘要

被引文献

相似文献

我们考虑了一类一般的正则化方法,它在线性测量的基础上学习参数向量。众所周知,如果正则化器是L2范数的非递减函数,那么学习到的向量就是输入数据的线性组合。这个结果被称为表征定理,是机器学习中基于核的方法的基础。本文在可微正则子的情况下,证明了上述条件的必要性。我们进一步将我们的分析扩展到学习矩阵的正则化方法,这个问题是由多任务学习的应用所激发的。在这种情况下,我们研究了一个更一般的表征定理,它适用于更大的正则化器类。给出了刻画这类矩阵正则器的一个充分必要条件,并给出了一些具有实际意义的具体例子。我们的分析使用了矩阵理论的基本原理,特别是矩阵非降函数的有用概念。
We consider a general class of regularization methods which learn a vector of parameters on the basis of linear measurements. It is well known that if the regularizer is a nondecreasing function of the L2 norm, then the learned vector is a linear combination of the input data. This result, known as the representer theorem, lies at the basis of kernel-based methods in machine learning. In this paper, we prove the necessity of the above condition, in the case of differentiable regularizers. We further extend our analysis to regularization methods which learn a matrix, a problem which is motivated by the application to multi-task learning. In this context, we study a more general representer theorem, which holds for a larger class of regularizers. We provide a necessary and sufficient condition characterizing this class of matrix regularizers and we highlight some concrete examples of practical importance. Our analysis uses basic principles from matrix theory, especially the useful notion of matrix nondecreasing functions.