Autonomous cleaning of corrupted scanned documents — A generative modeling approach

Autonomous cleaning of corrupted scanned documents — A generative modeling approach
复制标题

DOI:
10.1109/tpami.2014.2313126
复制
发表时间:
2012-01
期刊:
2012 IEEE Conference on Computer Vision and Pattern Recognition
影响因子:
--
通讯作者:
Zhenwen Dai;Jörg Lücke
Zhenwen Dai;Jörg Lücke
中科院分区:
其他
文献类型:
--
作者:
Zhenwen Dai;Jörg Lücke

文献摘要

被引文献

相似文献

我们研究清理被灰尘严重污染的扫描文本文档的任务,如手动线条笔划、溢出墨水等。我们的目标是仅基于页面包含的信息自动从单个字母大小的页面中清除灰尘。因此,我们的方法必须在没有监督的情况下学习字符表示法,并需要一种机制来区分学习的表示法和不规则的模式。为了学习字符表示,我们使用一个概率生成模型来参数化图案特征、特征方差、特征的平面排列和图案频率。模型的潜在变量描述了模式类别、模式位置以及单个模式特征的存在或不存在。利用一种新的变分EM近似对模型参数进行了优化。在学习之后,参数表示平面特征排列及其方差,而与其绝对位置无关。然后,基于所学习的表示定义的质量度量允许在构成污垢的规则字符图案和不规则图案之间进行自主区分。因此,可以去除不规则图案以清洁文档。对于完整的拉丁字母表,我们发现单个页面包含的字符示例不够多。然而,即使受到污垢的严重破坏,我们也表明,仅根据页面中包含的字符的结构规则性,可以高效地自主清理包含较少字符类型的页面。在使用来自不同字母表的字符的不同例子中,我们展示了该方法的一般性,并讨论了其对未来发展的影响。
We study the task of cleaning scanned text documents that are strongly corrupted by dirt such as manual line strokes, spilled ink etc. We aim at autonomously removing dirt from a single letter-size page based only on the information the page contains. Our approach, therefore, has to learn character representations without supervision and requires a mechanism to distinguish learned representations from irregular patterns. To learn character representations, we use a probabilistic generative model parameterizing pattern features, feature variances, the features' planar arrangements, and pattern frequencies. The latent variables of the model describe pattern class, pattern position, and the presence or absence of individual pattern features. The model parameters are optimized using a novel variational EM approximation. After learning, the parameters represent, independently of their absolute position, planar feature arrangements and their variances. A quality measure defined based on the learned representation then allows for an autonomous discrimination between regular character patterns and the irregular patterns making up the dirt. The irregular patterns can thus be removed to clean the document. For a full Latin alphabet we found that a single page does not contain sufficiently many character examples. However, even if heavily corrupted by dirt, we show that a page containing a lower number of character types can efficiently and autonomously be cleaned solely based on the structural regularity of the characters it contains. In different examples using characters from different alphabets, we demonstrate generality of the approach and discuss its implications for future developments.