HADARA – A Software System for Semi-Automatic Processing of Historical Handwritten Arabic Documents

HADARA – A Software System for Semi-Automatic Processing of Historical Handwritten Arabic Documents
复制标题

HADARA – 用于半自动处理历史手写阿拉伯语文档的软件系统

DOI:
10.2352/issn.2168-3204.2013.10.1.art00036
复制
发表时间:
2013
期刊:
Archiving Conference
影响因子:
--
通讯作者:
Jihad El
Jihad El
中科院分区:
--
文献类型:
--
作者:
Werner Pantke;Daniel Fecker;T. Fingscheidt;Abedelkadir Asi;Ofer Biller;Jihad El

文献摘要

被引文献

相似文献

最近,世界各地的许多大图书馆都在扫描他们的馆藏,使其公开提供,并保存历史文献。我们提出了一个模块化的软件系统,它可以作为一个工具,用于半自动处理的历史手写阿拉伯文文件。该系统的开发是HADARA项目的一部分,该项目旨在对阿拉伯语手稿进行历史文献分析,由一个项目小组组成,其中包括工程师和计算机科学家,但也包括语言学家和历史学家等用户。HADARA系统旨在支持对历史阿拉伯文文档的脚本和内容分析、识别和分类。该系统是按照迭代开发方法创建的,当前版本以交互方式和部分自动方式协助用户。在本文中,给出了一个系统的概述,并提出了第一个模块,支持在一个半自动的方式扫描手稿的注释。它们包括页面布局分析,文本行分割和转录。单词识别是在HADARA系统中实现的第一个应用程序,本文概述了它的概念。
Recently, many big libraries all over the world have been scanning their collections to make them publicly available and to preserve historical documents. We present a modular software system which can be used as a tool for semi-automatical processing of historical handwritten Arabic documents. The development of this system is part of the HADARA project which aims for historical document analysis of Arabic manuscripts and consists of a project team including engineers and computer scientists but also users such as linguists and historians. The HADARA system is designed to support script and content analysis, identification, and classification of historical Arabic documents. The system has been created following an iterative development approach, and the current version assists the user in an interactive and partially already in an automatic manner. In this paper, a system overview is given and the first modules are presented which support the annotation of a scanned manuscript in a semi-automatic manner. They comprise page layout analysis, text line segmentation, and transcription. Word spotting is the first application implemented in the HADARA system and its concept is outlined in this paper.