课题基金 / 基金详情

Development of a Repository for OCR Models and an Automatic Font Recognition tool OCR-D

Development of a Repository for OCR Models and an Automatic Font Recognition tool OCR-D
OCR 模型存储库和自动字体识别工具 OCR-D 的开发
批准号:
394448308
负责人:
Professor Dr. Manuel Burghardt, since 11/2019
金额:
$0.0万
依托单位:
依托单位国家:
德国
项目类别:
Research data and software (Scientific Library Services and Information Systems)
财政年份:
2018
资助国家:
德国
项目状态:
已结题
起止时间:
2017-12-31 至 2019-12-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
该项目解决了16至18世纪历史印刷品OCR识别率剧烈波动的问题,限制了VD16, VD17和VD18程序创建的全文数字化材料。在现代语料库上训练的识别模型缺乏历史印刷品或历史材料的细节,没有进行彻底的书目分析,与现代印刷品扫描通常达到的准确性相比,识别率会降低。在手动标记的基础上创建特定于字体的语料库是不现实的,因为印刷历史的重要知识是必要的,而且这种方法的可伸缩性是不够的。由于任务的重复性,这种方法也很容易出错。该项目将使人文学科能够在有限的努力下以特定字体的方式使用OCR。为了实现这一目标,该项目有三个主要目标:开发一个在线培训基础设施,允许为这些字体组和不同的OCR软件训练特定的模型。历史印刷品数字化中字体自动识别工具的开发。在这种情况下,首先使用在Typenrepertorium der Wiegendrucke中找到的ground truth来训练在incunabula中识别字体的算法。第二步,根据相似度对字体进行分组,以便在保持OCR准确性的同时获得尽可能少的组。提供模型存储库,其中开发的特定于字体的OCR模型可供公众使用。
英文摘要
The project addresses the problem of strongly fluctuating recognition rates of OCR for 16th to 18th century historical prints, limiting the full-text digitization of material created by the VD16, VD17, and VD18 programs.Recognition models trained on modern corpora lacking the specifics of historical prints or historic material without thorough bibliographic analysis, retard recognition rates in comparison to the accuracy now routinely achieved for scans of modern prints.The creation of font-specific corpora on the basis of manual tagging is unrealistic, since both non-trivial knowledge of printing history is necessary and the scalability of such an approach would be insufficient. Due to the repetitiveness of the task, such an approach is also very error-prone. The project will enable the humanities to use OCR in a font-specific manner with limited effort. In order to achieve this the project has three main objectives:The development of an online training infrastructure that allows specific models to be trained for these font groups and at the same time for different OCR software.Development of a tool for the automatic recognition of fonts in digitizations of historical prints. In this case, an algorithm for the recognition of fonts in incunabula is first trained using the ground truth found in the Typenrepertorium der Wiegendrucke. In a second step the fonts are grouped according to their similarity in order to get as few groups as possible while maintaining OCR accuracy.Provision of a model repository, in which developed font-specific OCR models are made available to the public.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金