On the Improvement of Recognizing Single-Line Strings of Japanese Historical Cursive

On the Improvement of Recognizing Single-Line Strings of Japanese Historical Cursive
复制标题

DOI:
10.1109/icdar.2019.00105
复制
发表时间:
2019-09
期刊:
2019 International Conference on Document Analysis and Recognition (ICDAR)
影响因子:
--
通讯作者:
A. Nagai
A. Nagai
中科院分区:
其他
文献类型:
--
作者:
A. Nagai

文献摘要

相似文献

对日本历史文献进行转录是将其作为文化资产保存的第一步。这些历史文献不仅可以直接用于防灾,还可以丰富日本文化。事实上,甚至有一个100多年来一直在进行的国家项目,目的是全面汇编旧文件。然而,即使是没有受过训练的现代日本人,也很难读懂日本的历史草书。本文以日本历史草书为研究对象,对其中的单行文字进行识别。我们的结果有超过95%的准确率,只有46个平假名字符组成的文本,而84.08%的准确率,包括数千个汉字字符的文本。两者都是最先进的精确度。也就是说,我们对46个平假名字符的识别结果明显优于以前的最先进水平。这是第一个识别数千个草书汉字的研究。此外,我们进行了各种实验来提高历史草书的识别准确率,包括稀有字符的数据增强,语言模型的增强,以及与测试数据相同作者编写的样本进行微调。因此,由于笔迹风格多种多样,实际上使用由同一作者编写的样本作为测试数据进行微调是有效的。它很容易超过数据增强和语言模型的改进。
Transcribing historical Japanese document is the first step to preserve them as cultural assets. These historical documents can be directly useful not only for disaster prevention but also for enriching Japanese culture. Indeed, there is even an ongoing national project for more than 100 years with the aim of comprehensively compiling old documents. However, it is difficult to read Japanese historical cursive even for modern Japanese without training. In this paper, we report on research to recognize a single-line text of Japanese historical cursive. Our result has more than 95% accuracy for the text consisting of only 46 Hiragana characters, while 84.08% accuracy for the text including thousands of Kanji characters. Both of them are state-of-the-art accuracy. That is, our result on 46 Hiragana characters significantly outperformed the previous state-of-the-art. And this is the first research to recognize thousands of cursive Kanji characters. Furthermore, we had various experiments to improve the recognition accuracy of historical cursive, which includes data augmentation for rare characters, enhancement by language model, and fine-tuning with samples written by the same author as the test data. As a result, because of various handwriting styles, it is practically effective to fine-tune with samples written by the same author as the test data. It easily outperforms the improvements by data augmentation and language model.