Text Retrieval for Japanese Historical Documents by Image Generation

Text Retrieval for Japanese Historical Documents by Image Generation
复制标题

DOI:
10.1145/3151509.3151512
复制
发表时间:
2017-11
期刊:
Proceedings of the 4th International Workshop on Historical Document Imaging and Processing
影响因子:
--
通讯作者:
Chisato Sugawara;Tomo Miyazaki;Yoshihiro Sugaya;S. Omachi
Chisato Sugawara;Tomo Miyazaki;Yoshihiro Sugaya;S. Omachi
中科院分区:
其他
文献类型:
--
作者:
Chisato Sugawara;Tomo Miyazaki;Yoshihiro Sugaya;S. Omachi

文献摘要

相似文献

历史文献数字化发展迅速。由于历史文献图像的数据量很大,因此文本检索是方便使用历史文献图像的一项重要技术。本文提出了一种基于文本查询的日语历史文献关键词检索方法。该方法自动生成查询文本的图像,并通过特征匹配在文档中检索与生成图像相似的区域。我们利用深度学习技术来生成接近日本历史文献图像文本的图像。此外,我们使用卷积神经网络提取对文档中文本的外观变化具有鲁棒性的特征,如文本的阴影和形状。我们在江户时代日本历史文献的公共数据集上进行了文本检索实验。实验结果表明了该方法的有效性。
Digitization of historical documents is growing rapidly. Text retrieval is a vital technology to facilitate the use of historical document images because of the large amount of data. In this paper, we propose a method for retrieving keywords in Japanese historical documents with text query. The proposed method automatically generates an image of the query text and retrieves regions in documents similar to the generated image by feature matching. We exploit a technique of deep learning to generate an image close to texts in Japanese historical document images. Furthermore, we use convolutional neural network to extract features robust to appearance variation of texts in documents, such as shade and shape of texts. We conducted the text retrieval experiments on the public dataset of Japanese historical documents in the Edo era. The experimental results show the effectiveness of the proposed method.