The nature of statistical learning theory~.

The nature of statistical learning theory~.
复制标题

DOI:
10.1109/tnn.1997.641482
复制
发表时间:
1997-01-01
影响因子:
--
通讯作者:
Cherkassky, V
Cherkassky, V
中科院分区:
其他
文献类型:
--
作者:
Cherkassky, V

文献摘要

被引文献

相似文献

统计学习理论(又名Vapnik-Chervonenkis或VC理论)最近出现了一个通用的数学框架,用于从有限样本中估计(学习)依赖关系。该理论结合了与学习相关的基本概念和原则,定义明确的问题制定和自洽的数学理论。因此,它(可以说)是预测学习的最佳理论,并且与其他通常基于直觉、渐近和/或生物学论点的更经验的方法相比,它是有利的。不幸的是,也许是因为它的数学严谨性和复杂性,这个理论在主流统计学和神经网络领域并没有得到很好的理解。VC理论有效地描述了小(有限)样本的统计估计。因此,该理论包括(作为特例)经典统计方法(为大样本和/或严格的参数假设而开发)。Vapnik的理论明确地考虑了样本大小,并提供了模型复杂性和可用训练数据之间的权衡的定量(分析)描述。该理论由四部分组成:1)经验风险最小化(ERM)归纳原则的一致性条件; 2)基于这些条件的学习机泛化能力的界限; 3)基于这些界限的小样本归纳推理原则; 4)实现上述归纳原则的构造性方法。尽管实践者最终对建设性的学习方法感兴趣,但根据Vapnik的说法,“没有什么比一个好的理论更实用了。事实上,理解统计学习理论对于理解构造性方法是必要的,因为学习理论的每一部分都是基于前一部分的。Vapnik的书只有187页,它没有数学证明,只有概念和主要结果的描述。但是,这不是一本供休闲阅读的书。这位评论家花了相当大的毅力(在六个月的时间里三次翻阅这本书)才完全理解VC理论的第1)-3)部分,这是理解第4)部分的建设性方法所必需的。原因是这本书描述了一个真正的数学理论,其中每个概念都有明确的含义,后面章节的结果是建立在以前的概念和结果之上的。本书的大部分内容描述了学习理论和概念。然而,最后一章描述了一种新的强大的构造性学习方法,称为支持向量机(SVM)。支持向量机方法是基于VC理论,最近的几项实证研究表明,它的成功,
Statistical Learning Theory (aka Vapnik–Chervonenkis or VC-theory) has recently emerged as a general mathematical framework for estimating (learning) dependencies from finite samples. This theory combines fundamental concepts and principles related to learning, well-defined problem formulation, and self-consistent mathematical theory. Thus, it is (arguably) the best available theory for predictive learning, and it compares favorably to other more empirical methodologies that are often based on intuitive, asymptotic, and/or biological arguments. Unfortunately, perhaps because of its mathematical rigor and complexity, this theory is not well understood in the mainstream statistics and in the field of neural networks. VC-theory effectively describes statistical estimation with small (finite) samples. Hence, this theory includes (as a special case) classical statistical methods (developed for large samples and/or strict parametric assumptions). Vapnik’s theory explicitly takes into account the sample size and provides quantitative (analytical) description of the tradeoff between the model complexity and the available training data. This theory consists of four parts: 1) conditions for consistency of the empirical risk minimization (ERM) inductive principle; 2) bounds on the generalization ability of learning machines based on these conditions; 3) principles for inductive inference from small samples based on these bounds; and 4) constructive methods for implementing above inductive principles. Whereas a practitioner is ultimately interested in constructive learning methods, according to Vapnik,“nothing is more practical than a good theory.” In fact, understanding statistical learning theory is necessary for understanding constructive methods, simply because each part of the learning theory is based on the preceding one. Vapnik’s book contains only 187 pages, and it has no mathematical proofs, only descriptions of concepts and main results. However, this is not a book for casual reading. It took this reviewer considerable persistence (three passes through the book over a six-month period) in order to fully understand parts 1)–3) of VC-theory that are necessary for understanding constructive methods in part 4). The reason is that the book describes a truly mathematical theory, where each concept has well-defined meaning, and the results in later chapters are built upon earlier concepts and results. Much of the book describes the learning theory and concepts. However, the last chapter describes a new powerful constructive learning methodology called support vector machines (SVM’s). The SVM method is based on VC-theory, and several recent empirical studies indicate its su-