课题基金 / 基金详情

Training of machine-learning based procedures for automated postcorrection of OCRed historical printings

Training of machine-learning based procedures for automated postcorrection of OCRed historical printings
基于机器学习的程序培训,用于 ORed 历史打印的自动后期校正
批准号:
431091758
负责人:
Professor Dr. Klaus U. Schulz
金额:
$0.0万
依托单位国家:
德国
项目类别:
Research Grants
财政年份:
2020
资助国家:
德国
项目状态:
已结题
起止时间:
2019-12-31 至 2021-12-31

项目摘要

项目成果

Professor Dr. Klaus U. Schulz的其他基金

相似基金

相关文献

中文摘要
翻译
历史印刷的OCR结果通常包含许多识别错误。因此,后校正方法在这一领域发挥着重要作用。一些自动化的后校正系统“单独”开发的一个特定的历史OCR语料库已经显示出良好的效果。然而,开发一个“全能”的通用系统,用于OCR结果的自动后校正,为不同的OCR引擎和任意的历史打印提供良好的结果,是一个雄心勃勃的未来目标。在OCR-D倡议的框架内,目前正在开发基于监督机器学习的OCR后校正系统。在理想的情况下,这些系统应该适用于任意OCR引擎和历史文本。在这个项目中,我们希望系统地研究训练数据和方法对校正结果质量的影响。长期的最终目标是发展一个“全能”(s.a.)后校正模型作为第一步,我们寻找训练数据和特征系统,为特定的OCR引擎和历史印刷类带来最佳的校正结果,分析其他OCR和语料库出现的校正问题。使用这些结果作为一个起点,我们寻找的方法,以尽量减少所需的额外努力(在地面真相准备和后训练方面)为更大和不均匀的语料库开发校正模型。要调查的具体点,除其他外,postcorrection模型的组合和自动选择的校正模型为一个给定的新的OCR语料库。
英文摘要
OCR-results for historical printings typically contain many recognition errors. Hence postcorrection methods play an important role in this field. Some automated postcorrection systems ``individually'' developed for a particular historical OCR-corpus have shown good results. However, the development of an ``omnipotent'' general system for automated postcorrection of OCR-results, offering good results for distinct OCR engines and arbitrary historical printings, is an ambitious future goal. In the framework of the OCR-D initiative currently OCR postcorrection systems are being developed that are based on supervised machine learning. In the ideal case these systems should be applicable to arbitrary OCR engines and historical texts. In this project we want to systematically study the influence of training data and -methods on the quality of the correction results achieved. The long-term ultimate goal is the development of an ``omnipotent'' (s.a.) postcorrection model. As a first step we look for training data and feature systems that lead to optimal correction results for specific OCR engines and classes of historical printings, analyzing correction problems arising for other OCRs and corpora. Using these results as a starting point we search for methods to minimize the additional effort needed (in terms of ground truth preparation and posttraining) for developing correction models for larger and inhomogeneous corpora. Specific points to be investigated are, among others, the combination of postcorrection models and the automated selection of a correction model for a given new OCR corpus.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Automated postcorrection of OCRed historical printings with integrated optional interactive postcorrection
  • 批准号:
    393215159
  • 项目类别:
    Research data and software (Scientific Library Services and Information Systems)
  • 资助金额:
    $0.0万
  • 财政年份:
    2018
  • 负责人:
    Professor Dr. Klaus U. Schulz
  • 依托单位:
Development of a web-based system for the postcorrection of historical OCR'ed texts
  • 批准号:
    314731081
  • 项目类别:
    Research data and software (Scientific Library Services and Information Systems)
  • 资助金额:
    $0.0万
  • 财政年份:
    2016
  • 负责人:
    Professor Dr. Klaus U. Schulz
  • 依托单位:
Domänen- und dokumentenadaptive Verfahren zur Nachkorrektur von OCR-Ergebnissen
Erweiterung eines Abfragemodells für XML-Daten zur interaktiven Exploration
国内基金
海外基金
Understanding structural evolution of galaxies with machine learning
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    Nicola Rosario Napolitano
  • 依托单位:
非标准随机调度模型的最优动态策略
  • 批准号:
    71071056
  • 项目类别:
    面上项目
  • 资助金额:
    28.0万元
  • 批准年份:
    2010
  • 负责人:
    吴贤毅
  • 依托单位:
微生物发酵过程的自组织建模与优化控制
  • 批准号:
    60704036
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    21.0万元
  • 批准年份:
    2007
  • 负责人:
    高学金
  • 依托单位: