The nature of statistical learning theory~.
The nature of statistical learning theory~.
复制标题
DOI:
10.1109/tnn.1997.641482
复制
发表时间:
1997-01-01
影响因子:
--
通讯作者:
Cherkassky, V
中科院分区:
文献类型:
--
作者:
Cherkassky, V
Statistical Learning Theory (aka Vapnik–Chervonenkis or VC-theory) has recently emerged as a general mathematical framework for estimating (learning) dependencies from finite samples. This theory combines fundamental concepts and principles related to learning, well-defined problem formulation, and self-consistent mathematical theory. Thus, it is (arguably) the best available theory for predictive learning, and it compares favorably to other more empirical methodologies that are often based on intuitive, asymptotic, and/or biological arguments. Unfortunately, perhaps because of its mathematical rigor and complexity, this theory is not well understood in the mainstream statistics and in the field of neural networks. VC-theory effectively describes statistical estimation with small (finite) samples. Hence, this theory includes (as a special case) classical statistical methods (developed for large samples and/or strict parametric assumptions). Vapnik’s theory explicitly takes into account the sample size and provides quantitative (analytical) description of the tradeoff between the model complexity and the available training data. This theory consists of four parts: 1) conditions for consistency of the empirical risk minimization (ERM) inductive principle; 2) bounds on the generalization ability of learning machines based on these conditions; 3) principles for inductive inference from small samples based on these bounds; and 4) constructive methods for implementing above inductive principles. Whereas a practitioner is ultimately interested in constructive learning methods, according to Vapnik,“nothing is more practical than a good theory.” In fact, understanding statistical learning theory is necessary for understanding constructive methods, simply because each part of the learning theory is based on the preceding one. Vapnik’s book contains only 187 pages, and it has no mathematical proofs, only descriptions of concepts and main results. However, this is not a book for casual reading. It took this reviewer considerable persistence (three passes through the book over a six-month period) in order to fully understand parts 1)–3) of VC-theory that are necessary for understanding constructive methods in part 4). The reason is that the book describes a truly mathematical theory, where each concept has well-defined meaning, and the results in later chapters are built upon earlier concepts and results. Much of the book describes the learning theory and concepts. However, the last chapter describes a new powerful constructive learning methodology called support vector machines (SVM’s). The SVM method is based on VC-theory, and several recent empirical studies indicate its su-