A Binary Convolutional Encoder-decoder Network for Real-time Natural Scene Text Processing

A Binary Convolutional Encoder-decoder Network for Real-time Natural Scene Text Processing
复制标题

DOI:
--
复制
发表时间:
2016-12
期刊:
ArXiv
影响因子:
--
通讯作者:
Zichuan Liu;Yixing Li;Fengbo Ren;Hao Yu
Zichuan Liu;Yixing Li;Fengbo Ren;Hao Yu
中科院分区:
其他
文献类型:
--
作者:
Zichuan Liu;Yixing Li;Fengbo Ren;Hao Yu

文献摘要

被引文献

相似文献

在本文中,我们开发了一个用于自然场景文本处理(NSTP)的二进制卷积编码器-解码器网络(B-CEDNet)。它将文本图像转换为一个类别区分的显著图,揭示了字符的类别,空间和形态信息。现有的解决方案要么消耗内存,要么消耗运行时间,无法应用于资源受限设备上的实时应用程序,如高级驾驶员辅助系统。所开发的网络可以通过一次性的前向操作处理包含字符的多个区域,并被训练为具有二进制权重和二进制特征图,这导致了显着的推理运行时间加速和内存使用减少。通过对超过20万幅合成场景文本图像(大小为32\times128 $)的训练,该算法在ICDAR-03和ICDAR-13数据集上的像素级精度分别达到90%和91%.它在GPU上实现的推理运行时间仅为4.59 ms,网络大小为2.14 MB,比全精度版本快8倍,节省96%。
In this paper, we develop a binary convolutional encoder-decoder network (B-CEDNet) for natural scene text processing (NSTP). It converts a text image to a class-distinguished salience map that reveals the categorical, spatial and morphological information of characters. The existing solutions are either memory consuming or run-time consuming that cannot be applied to real-time applications on resource-constrained devices such as advanced driver assistance systems. The developed network can process multiple regions containing characters by one-off forward operation, and is trained to have binary weights and binary feature maps, which lead to both remarkable inference run-time speedup and memory usage reduction. By training with over 200, 000 synthesis scene text images (size of $32\times128$), it can achieve $90\%$ and $91\%$ pixel-wise accuracy on ICDAR-03 and ICDAR-13 datasets. It only consumes $4.59\ ms$ inference run-time realized on GPU with a small network size of 2.14 MB, which is up to $8\times$ faster and $96\%$ smaller than it full-precision version.