The Early Japanese Books Text Line Segmentation base on Image Processing and Deep Learning

The Early Japanese Books Text Line Segmentation base on Image Processing and Deep Learning
复制标题

DOI:
10.1109/icamechs.2019.8861597
复制
发表时间:
2019-08
期刊:
2019 International Conference on Advanced Mechatronic Systems (ICAMechS)
影响因子:
--
通讯作者:
Bing Lyu;R. Akama;Hiroyuki Tomiyama;Lin Meng
Bing Lyu;R. Akama;Hiroyuki Tomiyama;Lin Meng
中科院分区:
其他
文献类型:
--
作者:
Bing Lyu;R. Akama;Hiroyuki Tomiyama;Lin Meng

文献摘要

被引文献

相似文献

早期的书籍记录了当时的政治、经济、文化、历史等大量信息。了解早期书籍可能有助于我们了解历史,成为近年来的一个重要研究课题。然而,在日本,很多早期的书籍都是用苦石字来描述的,现在没有使用过,只有少数专家能理解它。因此,大量的早期日语书籍仍然不被理解。目前,研究人员正试图借助图像处理和深度学习等计算机技术来理解早期的日语书籍。然而,这些书是由文章和图片组成的,有时会有在同一页上。这种情况增加了字符识别的难度,使得研究人员不得不分割文本行,将文章和图片分开。本文的目的是通过深度学习从扫描的早期日语书籍图像中分割出文本行。为了达到更好的精度,本文还提出了一种利用投影轮廓去除帧噪声的图像处理方法。实验结果表明,该方法的准确率、召回率和F值分别达到了95.2%、98.3%和96.6%,证明了该方法的有效性。
Early books record a lot of information such as politics, economy, culture, history at that time. Understanding the early books may help us know the history, becomes an important research recently. However, in Japan, a lot of early books are described by Kuzushi character, which is not used now and only few of specialists can understand it. Hence a large number of early Japanese books still do not be understood. Currently, researchers are trying to understand the early Japanese books with the assist of computer such as image processing and deep learning. However, these books are composed of articles and pictures, sometimes there are in the same page. This case increases the difficult of character recognition, and lets researchers have to segment the text line and separate the articles and pictures previously. This paper aims to segment the text line from the image of scanned early Japanese book by deep learning. For achieving better accuracy, this paper also proposes an image processing method which uses the projection profile for deleting the frame noise. The experimental results show that Precision, Recall and F Value achieve 95.2%, 98.3% and 96.6% respectively, and prove the effectiveness of our method.