An Introduction to Support Vector Machines and Other Kernel-based Learning Methods: Kernel-Induced Feature Spaces

An Introduction to Support Vector Machines and Other Kernel-based Learning Methods: Kernel-Induced Feature Spaces
复制标题

DOI:
10.1017/cbo9780511801389.005
复制
发表时间:
2000-03
期刊:
--
影响因子:
--
通讯作者:
N. Cristianini;J. Shawe-Taylor
N. Cristianini;J. Shawe-Taylor
中科院分区:
其他
文献类型:
--
作者:
N. Cristianini;J. Shawe-Taylor

文献摘要

被引文献

相似文献

Minsky 和 ​​Papert 在 20 世纪 60 年代强调了线性学习机有限的计算能力。一般来说,复杂的现实世界应用需要比线性函数更具表现力的假设空间。看待这个问题的另一种方式是,目标概念通常不能表达为给定属性的简单线性组合,但通常需要利用数据的更抽象特征。提出了多层阈值线性函数作为该问题的解决方案,这种方法导致了多层神经网络和学习算法(例如用于训练此类系统的反向传播)的发展。核表示提供了一种替代解决方案,将数据投影到高维特征空间中,以提高第 2 章线性学习机的计算能力。在对偶表示中使用线性机使得隐式执行此步骤成为可能。正如第 2 章所述,训练示例从来不会显得孤立,而是始终以示例对之间的内积形式出现。在对偶表示中使用机器的优点源于这样的事实:在该表示中,可调参数的数量不依赖于所使用的属性的数量。通过用适当选择的“内核”函数替换内积,只要内核计算与两个输入相对应的特征向量的内积,就可以隐式地执行到高维特征空间的非线性映射,而无需增加可调参数的数量。 […]
The limited computational power of linear learning machines was highlighted in the 1960s by Minsky and Papert. In general, complex real-world applications require more expressive hypothesis spaces than linear functions. Another way of viewing this problem is that frequently the target concept cannot be expressed as a simple linear combination of the given attributes, but in general requires that more abstract features of the data be exploited. Multiple layers of thresholded linear functions were proposed as a solution to this problem, and this approach led to the development of multi-layer neural networks and learning algorithms such as back-propagation for training such systems. Kernel representations offer an alternative solution by projecting the data into a high dimensional feature space to increase the computational power of the linear learning machines of Chapter 2. The use of linear machines in the dual representation makes it possible to perform this step implicitly. As noted in Chapter 2, the training examples never appear isolated but always in the form of inner products between pairs of examples. The advantage of using the machines in the dual representation derives from the fact that in this representation the number of tunable parameters does not depend on the number of attributes being used. By replacing the inner product with an appropriately chosen ‘kernel’ function, one can implicitly perform a non-linear mapping to a high dimensional feature space without increasing the number of tunable parameters, provided the kernel computes the inner product of the feature vectors corresponding to the two inputs. […]