Attempts to recognize anomalously deformed Kana in Japanese historical documents

Attempts to recognize anomalously deformed Kana in Japanese historical documents
复制标题

尝试识别日本历史文献中异常变形的假名

DOI:
10.1145/3151509.3151514
复制
发表时间:
2017
期刊:
Proceedings of the 4th International Workshop on Historical Document Imaging and Processing
影响因子:
--
通讯作者:
M. Nakagawa
M. Nakagawa
中科院分区:
--
文献类型:
--
作者:
Hung Tuan Nguyen;N. Ly;K. Nguyen;C. Nguyen;M. Nakagawa

文献摘要

被引文献

相似文献

本文提出了三种不同任务的方法来识别日本历史文献中异常变形的假名,这是IEICE PRMU1 2017提出的异议。任务分为三个层次:单字识别、三个假名字符序列识别和无限制假名识别。我们对每个任务比较几种方法。对于第一级,我们评估了基于CNN的方法和基于BLSTM的方法。对于第2级,我们考虑了CNN和BLSTM组合架构的几种变体。对于第3级,我们比较了第2级方法的扩展和基于分割的方法。单个字符识别准确率为96.8%,三个假名字符序列识别准确率为87.12%,无限制假名识别准确率为73.3%。这些结果证明了CNN和BLSTM在这些任务上的性能。
This paper presents methods for three different tasks of recognizing anomalously deformed Kana in Japanese historical documents, which were contested by IEICE PRMU1 2017. The tasks have three levels: single character recognition, three Kana characters sequence recognition and unrestricted Kana recognition. We compare several methods for each task. For the level 1, we evaluate CNN based methods and BLSTM based methods. For the level 2, we consider several variations of a combined architecture of CNN and BLSTM. For the level 3, we compare an extension of the method for the level 2 and a segmentation based method. We achieve the single character recognition accuracy of 96.8%, the three Kana characters sequence recognition accuracy of 87.12% and the unrestricted Kana recognition accuracy of 73.3%. These results prove the performance of CNN and BLSTM on these tasks.