Virtual unrolling and information recovery from scanned scrolled historical documents

Virtual unrolling and information recovery from scanned scrolled historical documents
复制标题

DOI:
10.1016/j.patcog.2013.06.015
复制
发表时间:
2014-01-01
影响因子:
8
通讯作者:
Rosin, Paul L.
Rosin, Paul L.
中科院分区:
计算机科学1区
文献类型:
--
作者:
Samko, Oksana;Lai, Yu-Kun;Rosin, Paul L.

文献摘要

被引文献

相似文献

我们工作的目标是使阅读脆弱的滚动历史羊皮纸,而不需要物理解开他们,从而提供有价值的信息,以广泛的学术学科。由于需要羊皮纸扫描技术,这个问题还没有被计算机视觉界正确地研究:标准的X射线设备是不够的,因为除了羊皮纸的底层结构之外,还需要提取出羊皮纸墨水。有效的数据恢复也受到损害,因为由于羊皮纸的变质,历史滚动文档的内容无法访问。我们创建一个滚动的羊皮纸的基础几何的3D体积模型,并执行羊皮纸的数字展开,产生一个可读的文本图像作为输出。建议的恢复框架包括结构保持各向异性过滤结合强大的分割,表面建模和墨水投影。我们展示了真实的例子,我们的算法是如何能够恢复底层的文本,并解决滚动羊皮纸分析,即分割连接层和处理数据,而无需用户交互的主要挑战。(C)2013爱思唯尔有限公司保留所有权利。
The objective of our work is to enable the reading of fragile scrolled historical parchments without the need to physically unravel them, thus providing valuable information to a wide range of scholarly disciplines. This problem has not been investigated by the computer vision community properly yet due to the need for parchment scanning technology: standard X-ray equipment is not sufficient as there is a requirement to extract out parchment ink in addition to the parchment's underlying structure. Effective data recovery is also compromised as content from historical scrolled documents is inaccessible due to the deterioration of the parchment. We create a 3D volumetric model of a scrolled parchment's underlying geometry and perform digital unwrapping of the parchment, producing a readable image of the text as an output. The proposed recovery framework consists of structure preserving anisotropic filtering in combination with robust segmentation, surface modelling and ink projection. We demonstrate with real examples how our algorithm is able to recover the underlying text and to solve the major challenge for scrolled parchment analysis, namely segmentation of connected layers and processing the data without user interaction. (C) 2013 Elsevier Ltd. All rights reserved.