Recognizing characters of ancient manuscripts

Recognizing characters of ancient manuscripts
复制标题

识别古代手稿的字符

DOI:
--
复制
发表时间:
2010
期刊:
Electronic imaging
影响因子:
--
通讯作者:
Robert Sablatnig
Robert Sablatnig
中科院分区:
--
文献类型:
--
作者:
Markus Diem;Robert Sablatnig

文献摘要

被引文献

相似文献

考虑到印刷拉丁文本,光学字符识别 (OCR) 系统的主要问题已得到解决。然而,对于降级的手写文档图像,二值化等基本预处理步骤使用最先进的方法获得的结果很差。本文对 11 世纪的古代斯拉夫手稿进行了调查。为了最大限度地减少错误字符分割的后果,提出了一种基于局部描述符的免二值化方法。此外,本地信息还可以识别部分可见或褪色的字符。所提出的算法包括两个步骤:字符分类和字符定位。最初提取尺度不变特征变换(SIFT)特征,随后使用支持向量机(SVM)对其进行分类。然后,根据兴趣点的空间信息对兴趣点进行聚类。因此,基于预先分类的局部描述符的加权投票方案,字符被定位并最终被识别。初步结果表明,所提出的系统可以处理具有背景杂乱(例如污渍、撕裂)和淡出字符的高度退化的手稿图像。
Considering printed Latin text, the main issues of Optical Character Recognition (OCR) systems are solved. However, for degraded handwritten document images, basic preprocessing steps such as binarization, gain poor results with state-of-the-art methods. In this paper ancient Slavonic manuscripts from the 11th century are investigated. In order to minimize the consequences of false character segmentation, a binarization-free approach based on local descriptors is proposed. Additionally local information allows the recognition of partially visible or washed out characters. The proposed algorithm consists of two steps: character classification and character localization. Initially Scale Invariant Feature Transform (SIFT) features are extracted which are subsequently classified using Support Vector Machines (SVM). Afterwards, the interest points are clustered according to their spatial information. Thereby, characters are localized and finally recognized based on a weighted voting scheme of pre-classified local descriptors. Preliminary results show that the proposed system can handle highly degraded manuscript images with background clutter (e.g. stains, tears) and faded out characters.