A Comparative Study of Two State-of-the-Art Feature Selection Algorithms for Texture-Based Pixel-Labeling Task of Ancient Documents

A Comparative Study of Two State-of-the-Art Feature Selection Algorithms for Texture-Based Pixel-Labeling Task of Ancient Documents
复制标题

古代文献基于纹理的像素标记任务的两种最先进特征选择算法的比较研究

DOI:
--
复制
发表时间:
2018
期刊:
影响因子:
3.2
通讯作者:
N. Amara
N. Amara
中科院分区:
--
文献类型:
--
作者:
Maroua Mehri;Ramzi Chaieb;Karim Kalti;P. Héroux;R. Mullot;N. Amara

文献摘要

被引文献

相似文献

近年来,纹理特征被广泛应用于历史文献图像分析。然而,很少有研究专门关注用于历史文档图像分析的特征选择算法。事实上,已经出现了在数据挖掘和机器学习任务中使用特征选择算法的重要需求,因为它有助于降低数据维度并提高算法性能,例如像素分类算法。因此,本文在分析和选择纹理特征的基础上,采用一种经典的像素标注方法,对遗传算法和ReliefF算法这两种传统的特征选择算法进行了比较研究。本研究中的两种特征选择算法被应用于HBR数据集的一个训练集,以推导出每个分析的基于纹理的特征集中选择最多的纹理特征。评价的特征集包括许多最新的纹理特征(Tamura、局部二值模式、灰度游程矩阵、自相关函数、灰度共生矩阵、Gabor滤波器、三级Haar小波变换、三级Daubechies滤波器的三级小波变换和四级Daubechies滤波器的三级小波变换)。在我们的实验中,使用了在历史书籍识别比赛(HBR2013数据集:Prima,Salford,UK)的上下文中提供的公共历史文档图像语料库。本文给出了定性和数值实验,旨在根据所使用的纹理特征集,对每种评估的特征选择算法的优缺点提供一套全面的指导方针。
Recently, texture features have been widely used for historical document image analysis. However, few studies have focused exclusively on feature selection algorithms for historical document image analysis. Indeed, an important need has emerged to use a feature selection algorithm in data mining and machine learning tasks, since it helps to reduce the data dimensionality and to increase the algorithm performance such as a pixel classification algorithm. Therefore, in this paper we propose a comparative study of two conventional feature selection algorithms, genetic algorithm and ReliefF algorithm, using a classical pixel-labeling scheme based on analyzing and selecting texture features. The two assessed feature selection algorithms in this study have been applied on a training set of the HBR dataset in order to deduce the most selected texture features of each analyzed texture-based feature set. The evaluated feature sets in this study consist of numerous state-of-the-art texture features (Tamura, local binary patterns, gray-level run-length matrix, auto-correlation function, gray-level co-occurrence matrix, Gabor filters, Three-level Haar wavelet transform, three-level wavelet transform using 3-tap Daubechies filter and three-level wavelet transform using 4-tap Daubechies filter). In our experiments, a public corpus of historical document images provided in the context of the historical book recognition contest (HBR2013 dataset: PRImA, Salford, UK) has been used. Qualitative and numerical experiments are given in this study in order to provide a set of comprehensive guidelines on the strengths and the weaknesses of each assessed feature selection algorithm according to the used texture feature set.