Stroke-Based Scene Text Erasing Using Synthetic Data for Training

Stroke-Based Scene Text Erasing Using Synthetic Data for Training
复制标题

使用合成数据训练的基于笔划的场景文本擦除

DOI:
10.1109/tip.2021.3125260
复制
发表时间:
2021-04
影响因子:
10.6
通讯作者:
Zhengmi Tang;Tomo Miyazaki;Yoshihiro Sugaya;S. Omachi
Zhengmi Tang;Tomo Miyazaki;Yoshihiro Sugaya;S. Omachi
中科院分区:
计算机科学1区
文献类型:
--
作者:
Zhengmi Tang;Tomo Miyazaki;Yoshihiro Sugaya;S. Omachi

文献摘要

相似文献

场景文本擦除是将自然图像中的文本区域替换为合理的内容,近年来引起了计算机视觉界的极大关注。在场景文本擦除中有两个潜在的子任务:文本检测和图像修复。这两个子任务都需要大量的数据来实现更好的性能;然而,缺乏大规模的真实世界场景文本删除数据集并不允许现有的方法实现其潜力。为了弥补成对真实世界数据的不足,我们在额外增强后大量使用了合成文本,随后仅在改进的合成文本引擎生成的数据集上训练我们的模型。我们提出的网络包含一个笔划掩码预测模块和背景修复模块,可以从裁剪的文本图像中提取文本笔划作为一个相对较小的洞,以保持更多的背景内容,以获得更好的修复效果。该模型可以部分擦除场景图像中的文本实例与边界框或与现有的场景文本检测器自动场景文本擦除。对SCUT-Syn、ICDAR 2013和SCUT-EnsText数据集进行定性和定量评估的实验结果表明,即使在真实数据上进行训练,我们的方法也明显优于现有的最先进方法。
Scene text erasing, which replaces text regions with reasonable content in natural images, has drawn significant attention in the computer vision community in recent years. There are two potential subtasks in scene text erasing: text detection and image inpainting. Both subtasks require considerable data to achieve better performance; however, the lack of a large-scale real-world scene-text removal dataset does not allow existing methods to realize their potential. To compensate for the lack of pairwise real-world data, we made considerable use of synthetic text after additional enhancement and subsequently trained our model only on the dataset generated by the improved synthetic text engine. Our proposed network contains a stroke mask prediction module and background inpainting module that can extract the text stroke as a relatively small hole from the cropped text image to maintain more background content for better inpainting results. This model can partially erase text instances in a scene image with a bounding box or work with an existing scene-text detector for automatic scene text erasing. The experimental results from the qualitative and quantitative evaluation on the SCUT-Syn, ICDAR2013, and SCUT-EnsText datasets demonstrate that our method significantly outperforms existing state-of-the-art methods even when they are trained on real-world data.