Information Theoretic Learning

Information Theoretic Learning
复制标题

DOI:
10.4018/978-1-59904-849-9.ch133
复制
发表时间:
2005
期刊:
--
影响因子:
--
通讯作者:
Deniz Erdoğmuş;J. Príncipe
Deniz Erdoğmuş;J. Príncipe
中科院分区:
其他
文献类型:
--
作者:
Deniz Erdoğmuş;J. Príncipe

文献摘要

被引文献

相似文献

学习系统依赖于三个相互关联的组成部分:拓扑、成本/性能函数和学习算法。拓扑为映射提供了约束,学习算法提供了找到最优解的方法;但解对什么是最优的呢?最优性以准则为特征,在神经网络文献中,这是最少被提及的成分,但它对泛化性能有决定性的影响。当然,应该更好地理解和调查选择标准背后的假设。传统上,最小二乘一直是回归问题的基准准则;考虑到分类是一个估计类后验概率的回归问题,最小二乘法被用来训练神经网络和其他分类器拓扑来近似正确的标签。在回归中使用最小二乘的主要动机仅仅来自于这个标准提供的智力安慰,因为它在传统的线性最小二乘回归应用中取得了成功——它可以简化为求解线性方程组。对于非线性回归,可以强调测量误差的高斯性假设,并结合极大似然原理来推广该准则。在非参数回归中,最小二乘原理导致条件期望解,直观上吸引人。虽然这些都是使用均方误差作为成本的好理由,但它本质上与上述假设和习惯有关。因此,在非高斯分布条件下,当坚持二阶统计准则时,在非线性自适应系统的训练过程中,误差信号中存在未捕获的信息。这个论点延伸到其他线性二阶技术,如主成分分析(PCA)、线性判别分析(LDA)和典型相关分析(CCA)。最近的工作试图通过利用核技术或其他启发式方法将这些技术推广到非线性场景。这就引出了一个问题:还有什么其他的成本函数可以用来训练自适应系统?我们如何建立严格的技术,将有用的概念从线性和二阶统计技术扩展到非线性和高阶统计学习方法?
INTRODUCTION Learning systems depend on three interrelated components: topologies, cost/performance functions, and learning algorithms. Topologies provide the constraints for the mapping, and the learning algorithms offer the means to find an optimal solution; but the solution is optimal with respect to what? Optimality is characterized by the criterion and in neural network literature, this is the least addressed component, yet it has a decisive influence in generalization performance. Certainly, the assumptions behind the selection of a criterion should be better understood and investigated. Traditionally, least squares has been the benchmark criterion for regression problems; considering classification as a regression problem towards estimating class posterior probabilities, least squares has been employed to train neural network and other classifier topologies to approximate correct labels. The main motivation to utilize least squares in regression simply comes from the intellectual comfort this criterion provides due to its success in traditional linear least squares regression applications – which can be reduced to solving a system of linear equations. For nonlinear regression, the assumption of Gaussianity for the measurement error combined with the maximum likelihood principle could be emphasized to promote this criterion. In nonparametric regression, least squares principle leads to the conditional expectation solution, which is intuitively appealing. Although these are good reasons to use the mean squared error as the cost, it is inherently linked to the assumptions and habits stated above. Consequently, there is information in the error signal that is not captured during the training of nonlinear adaptive systems under non-Gaussian distribution conditions when one insists on secondorder statistical criteria. This argument extends to other linear-second-order techniques such as principal component analysis (PCA), linear discriminant analysis (LDA), and canonical correlation analysis (CCA). Recent work tries to generalize these techniques to nonlinear scenarios by utilizing kernel techniques or other heuristics. This begs the question: what other alternative cost functions could be used to train adaptive systems and how could we establish rigorous techniques for extending useful concepts from linear and second-order statistical techniques to nonlinear and higher-order statistical learning methodologies?