Adaptive document image binarization

Adaptive document image binarization
复制标题

DOI:
10.1016/s0031-3203(99)00055-2
复制
发表时间:
2000-02-01
影响因子:
8
通讯作者:
Pietikäinen, M
Pietikäinen, M
中科院分区:
计算机科学1区
文献类型:
--
作者:
Sauvola, J;Pietikäinen, M

文献摘要

被引文献

相似文献

提出了一种新的自适应文档图像二值化方法,其中页面被视为文本、背景和图片等子组件的集合。噪声,照明和许多源类型相关的退化所造成的问题得到解决。两个新的算法被应用到确定每个像素的局部阈值。该算法的性能评估利用测试图像与地面真理,评估指标的文本和合成图像的二值化,并基于权重的排名程序的最终结果呈现。所提出的算法进行了测试,包括不同类型的文档组件和退化的图像。结果与文献中的一些已知技术进行了比较。基准测试的结果表明,该方法适应和性能良好,在每种情况下,定性和定量。(C)1999年模式识别学会。由Elsevier Science Ltd.出版,版权所有。
A new method is presented for adaptive document image binarization, where the page is considered as a collection of subcomponents such as text, background and picture. The problems caused by noise, illumination and many source type-related degradations are addressed. Two new algorithms are applied to determine a local threshold for each pixel. The performance evaluation of the algorithm utilizes test images with ground-truth, evaluation metrics for binarization of textual and synthetic images, and a weight-based ranking procedure for the final result presentation. The proposed algorithms were tested with images including different types of document components and degradations. The results were compared with a number of known techniques in the literature. The benchmarking results show that the method adapts and performs well in each case qualitatively and quantitatively. (C) 1999 Pattern Recognition Society. Published by Elsevier Science Ltd. All rights reserved.