Toward intelligent training of supervised image classifications: directing training data acquisition for SVM classification

Toward intelligent training of supervised image classifications: directing training data acquisition for SVM classification
复制标题

DOI:
10.1016/j.rse.2004.06.017
复制
发表时间:
2004-10-30
影响因子:
13.5
通讯作者:
Mathur, A
Mathur, A
中科院分区:
工程技术1区
文献类型:
--
作者:
Foody, GM;Mathur, A

文献摘要

被引文献

相似文献

训练监督图像分类的传统方法旨在在频谱上完全描述所有类别。为了实现特征空间中每个类的完整描述,通常需要大的训练集。然而,并不总是需要训练统计数据来提供对类的完整和代表性的描述,特别是在使用非参数分类器的情况下。对于通过支持向量机的分类,仅需要是支持向量的训练样本,其位于特征空间中的类分布的边缘的一部分上;所有其他训练样本对分类分析没有贡献。如果可以在分类之前识别可能提供支持向量的区域,则可以智能地选择有用的训练样本。以有用的训练样本为目标的能力可以允许从小训练集进行准确分类。为从多谱段卫星传感器数据中对农作物进行分类,探索了这种智能训练样本收集的潜力。使用传统的训练方法,只有四分之一的训练样本实际上对分析做出了积极的贡献,并允许作物被分类到高准确度(92.5%)。因此,大多数训练集是不必要的,因为它对分析没有贡献。然而,利用土壤类型的辅助信息,有可能限制训练样本的获取过程。通过将训练样本采集仅限于特定土壤类型的区域,可以使用较小的训练集对数据进行分类,而不会损失准确性。因此,可以使用少量智能选择的训练样本来对数据集进行分类,其准确度与以常规方式导出的较大训练集一样。结果表明,有可能直接训练数据采集策略的目标是最有用的训练样本,以实现高效和准确的图像分类。(C)2004年爱思唯尔公司All rights reserved.
Conventional approaches to training a supervised image classification aim to fully describe all of the classes spectrally. To achieve a complete description of each class in feature space, a large training set is typically required. It is not, however, always necessary to have training statistics that provide a complete and representative description of the classes, especially if using nonparametric classifiers. For classification by a support vector machine, only the training samples that are support vectors, which lie on part of the edge of the class distribution in feature space, are required; all other training samples provide no contribution to the classification analysis. If regions likely to furnish support vectors can be identified in advance of the classification, it may be possible to intelligently select useful training samples. The ability to target useful training samples may allow accurate classification from small training sets. This potential for intelligent training sample collection was explored for the classification of agricultural crops from multispectral satellite sensor data. With a conventional approach to training, only a quarter of the training samples acquired actually made a positive contribution to the analysis and allowed the crops to be classified to a high accuracy (92.5%). The majority of the training set, therefore, was unnecessary as it made no contribution to the analysis. Using ancillary information on soil type, however, it would be possible to constrain the training sample acquisition process. By limiting training sample acquisition only to regions with a specific soil type, it was possible to use a small training set to classify the data without loss of accuracy. Thus, a small number of intelligently selected training samples may be used to classify a data set as accurately as a larger training set derived in a conventional manner. The results illustrate the potential to direct training data acquisition strategies to target the most useful training samples to allow efficient and accurate image classification. (C) 2004 Elsevier Inc. All rights reserved.