General Models for Handwritten Text Recognition: Feasibility and State-of-the Art. German Kurrent as an Example

General Models for Handwritten Text Recognition: Feasibility and State-of-the Art. German Kurrent as an Example
复制标题

手写文本识别的通用模型:可行性和最新技术。

DOI:
10.5334/johd.46
复制
发表时间:
2021
影响因子:
--
通讯作者:
Jake Purcell
Jake Purcell
中科院分区:
--
文献类型:
--
作者:
Tobias Hodel;David Schoch;Christa Schneider;Jake Purcell

文献摘要

被引文献

相似文献

现有的文本识别引擎能够训练通用模型来识别特定脚本内的不仅是一只特定的手,而且是来自相当长的时间段(超过100年)的多个历史手。本文比较了不同的文本识别引擎以及它们在独立于训练和验证集的测试集上的性能。我们认为,测试集和地面真理都应该由研究人员作为共享任务的一部分提供,以便对发动机进行比较。这将为需要认可模式的机构提供一系列可能的选择。作为测试集,我们提供了一个由2426行组成的数据集,这些行是从1848年至1903年瑞士联邦委员会的会议纪要中随机选择的。据我们所知,无论是我们认为是基本事实的上述文本行,还是该语料库中的众多不同的手,都从未被用于训练手写文本识别模型。此外,由于其可变性和所跨越的时间范围,所使用的数据集非常适合于进行涉及识别引擎和大型训练集的比较。因此,本文认为,两个测试引擎,HTR+和PYLAIA,都可以处理大型训练集。所得到的模型在由未知但风格相似的手组成的测试集上产生了非常好的结果。
Existing text recognition engines enables to train general models to recognize not only one specific hand but a multitude of historical hands within a particular script, and from a rather large time period (more than 100 years). This paper compares different text recognition engines and their performance on a test set independent of the training and validation sets. We argue that both, test set and ground truth, should be made available by researchers as part of a shared task to allow for the comparison of engines. This will give insight into the range of possible options for institutions in need of recognition models. As a test set, we provide a data set consisting of 2,426 lines which have been randomly selected from meeting minutes of the Swiss Federal Council from 1848 to 1903. To our knowledge, neither the aforementioned text lines, which we take as ground truth, nor the multitude of different hands within this corpus have ever been used to train handwritten text recognition models. In addition, the data set used is perfect for making comparisons involving recognition engines and large training sets due to its variability and the time frame it spans. Consequently, this paper argues that both the tested engines, HTR+ and PyLaia, can handle large training sets. The resulting models have yielded very good results on a test set consisting of unknown but stylistically similar hands.