Generation of a Large-Scale Line Image Dataset with Ground Truth Texts from Page-Level Autograph Documents

Generation of a Large-Scale Line Image Dataset with Ground Truth Texts from Page-Level Autograph Documents
复制标题

DOI:
10.1007/978-3-030-92185-9_29
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
A. Nagai
A. Nagai
中科院分区:
其他
文献类型:
--
作者:
A. Nagai

文献摘要

相似文献

近年来,深度学习技术对日本历史草书的识别具有很高的准确性。然而,大多数已知的草书数据集是从印刷文档中收集的,这些文档是为一般公众编写的,易于阅读。我们的研究旨在提高签名文件的识别,签名文件比打印文件更难识别,因为它们通常是私人的,并且以各种写作风格写成。为了创建有用的签名文档数据集,本文设计了一种技术,给定签名文档的GT转录仅在页面级别可用,该技术可以生成伴随相应的ground truth (GT)文本的许多行图像。我们的方法使用HRNet进行线检测,使用CRNN进行线识别。利用HRNet将页面图像分解为行,这些行映射到GT文本,而GT文本与GT转录分开分解,通过光束搜索解决基于相似性的对齐问题。我们引入了两个对齐思路:允许不相邻的线的乱序映射和允许多对多映射。通过这两个正交的想法,我们得到了一个由43,271个可靠的签名线图像映射到GT文本的数据集。通过将该数据集与打印数据集一起从头开始训练CRNN,提高了签名文档的识别精度。
Recently, Deep Learning techniques help to recognize Japanese historical cursive with high accuracy. However, most of the known cursive dataset have been gathered from printed documents which are written for the general public and easy to read. Our research aims to improve the recognition of autograph documents, which are more difficult to recognize than printed documents because they are often private and written in various writing styles. To create a useful autograph document dataset, this paper devises a technique to generate many line images accompanied by the corresponding ground truth (GT) texts, given an autograph document whose GT transcription is available only at the page-level.Our method utilizes HRNet for line detection and CRNN for line recognition. HRNet is used to decompose the page image into lines that is mapped to GT text, which is decomposed separately from GT transcription, by similarity-based alignment solved by beam search. We introduce two ideas to the alignment: to allow out-of-order mapping of the lines not adjacent to each other and to allow many-to-many mapping. With these orthogonal two ideas, we obtained a dataset consisting of 43,271 reliable autograph line images mapped to GT texts. By training CRNN from scratch on this dataset together with printed dataset, recognition accuracy for autograph documents is improved.