A dictionary learning and KPCA-based feature extraction method for off-line handwritten Tibetan character recognition

A dictionary learning and KPCA-based feature extraction method for off-line handwritten Tibetan character recognition
复制标题

DOI:
10.1016/j.ijleo.2015.07.144
复制
发表时间:
2015-12
期刊:
影响因子:
3.1
通讯作者:
Heming Huang;F. Da
Heming Huang;F. Da
中科院分区:
物理与天体物理3区
文献类型:
--
作者:
Heming Huang;F. Da

文献摘要

相似文献

核主成分分析(KPCA)是一种强大的特征提取技术。然而,对于字符识别,由于每个类别的样本量较大,KPCA 的计算成本太高。提出了一种基于字典学习和KPCA的新型两阶段特征提取方法DL-KPCA用于字符识别。第一阶段,采用字典学习方法K-SVD,首先从每一类的原始样本集中构造一个代表性样本子集。然后,对于测试样本,从所有构建的样本子集的并集中找到其K个最近邻,并将其最近邻的类作为候选类。第二阶段,利用KPCA将测试样本及其构建的候选类样本子集变换到特征空间,最后在特征空间中利用K-NN对测试样本进行分类。在新开发的藏文手写体字符样本数据库THCDB和改组后的USPS数字数据库上的实验结果表明,对于字符识别问题,用所提出的DL-KPCA提取特征是可行的。
Kernel principal component analysis (KPCA) is a powerful feature extraction technique. For character recognition, however, the computation cost of KPCA is too high because of much larger sample size of each class. A novel two-stage feature extraction method DL-KPCA that based on dictionary learning and KPCA is proposed for character recognition. In the first stage, with the dictionary learning method K-SVD, a representative sample subset is constructed from the original sample set of each class at first. Then, to the test sample, find its K nearest neighbors from the union of all the constructed sample subsets and consider the classes of their nearest neighbors as the candidate classes. In the second stage, the test sample and the constructed sample subsets of its candidate classes are transformed to the feature space with KPCA, and the test sample is finally classified with K-NN in the feature space. Experimental results on THCDB, a recently developed Tibetan handwritten character sample database, and the reshuffled USPS digit database show that, to character recognition problems, it is feasible to extract the features with the proposed DL-KPCA.