Intrinsic Decomposition of Document Images In-the-Wild

Intrinsic Decomposition of Document Images In-the-Wild
复制标题

DOI:
--
复制
发表时间:
2020-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Sagnik Das;H. Sial;Ke Ma;R. Baldrich;M. Vanrell;D. Samaras
Sagnik Das;H. Sial;Ke Ma;R. Baldrich;M. Vanrell;D. Samaras
中科院分区:
其他
文献类型:
--
作者:
Sagnik Das;H. Sial;Ke Ma;R. Baldrich;M. Vanrell;D. Samaras

文献摘要

相似文献

自动文档内容处理会受到纸张形状、照明条件颜色不均匀和多样化造成的伪影的影响。由于需要大量数据,对真实数据进行完全监督的方法是不可能的。因此,当前最先进的深度学习模型是在完全或部分合成图像上进行训练的。然而,文档阴影或阴影去除结果仍然受到影响,因为:(a)现有方法依赖于局部颜色统计的均匀性,这限制了它们在具有复杂文档形状和纹理的真实场景中的应用; (b) 使用具有非现实模拟照明条件的合成或混合数据集来训练模型。在本文中,我们通过两个主要贡献来解决这些问题。首先,一种基于物理约束的学习方法,根据固有图像形成直接估计文档反射率,该方法可推广到具有挑战性的照明条件。其次,一个新的数据集通过添加大量逼真的阴影和多样化的多光源条件,明显改进了以前的合成数据集,并且经过专门定制以处理野外文档。所提出的架构以自我监督的方式工作,其中仅将合成纹理用作弱训练信号(消除了对具有着色和反射率解缠结版本的非常昂贵的地面实况的需要)。所提出的方法使具有挑战性的照明的真实场景中的文档反射率估计得到了显着的推广。我们对可用于内在图像分解和文档阴影去除任务的真实基准数据集进行了广泛的评估。我们的反射率估计方案在用作 OCR 管道的预处理步骤时,字符错误率 (CER) 提高了 26%,从而证明了实际适用性。
Automatic document content processing is affected by artifacts caused by the shape of the paper, non-uniform and diverse color of lighting conditions. Fully-supervised methods on real data are impossible due to the large amount of data needed. Hence, the current state of the art deep learning models are trained on fully or partially synthetic images. However, document shadow or shading removal results still suffer because: (a) prior methods rely on uniformity of local color statistics, which limit their application on real-scenarios with complex document shapes and textures and; (b) synthetic or hybrid datasets with non-realistic, simulated lighting conditions are used to train the models. In this paper we tackle these problems with our two main contributions. First, a physically constrained learning-based method that directly estimates document reflectance based on intrinsic image formation which generalizes to challenging illumination conditions. Second, a new dataset that clearly improves previous synthetic ones, by adding a large range of realistic shading and diverse multi-illuminant conditions, uniquely customized to deal with documents in-the-wild. The proposed architecture works in a self-supervised manner where only the synthetic texture is used as a weak training signal (obviating the need for very costly ground truth with disentangled versions of shading and reflectance). The proposed approach leads to a significant generalization of document reflectance estimation in real scenes with challenging illumination. We extensively evaluate on the real benchmark datasets available for intrinsic image decomposition and document shadow removal tasks. Our reflectance estimation scheme, when used as a pre-processing step of an OCR pipeline, shows a 26% improvement of character error rate (CER), thus, proving the practical applicability.