Word spotting in Chinese document images without layout analysis

Word spotting in Chinese document images without layout analysis
复制标题

中文文档图像中的文字识别,无需布局分析

DOI:
--
复制
发表时间:
2002
期刊:
Object recognition supported by user interaction for service robots
影响因子:
--
通讯作者:
C. Tan
C. Tan
中科院分区:
--
文献类型:
--
作者:
Yue Lu;C. Tan

文献摘要

被引文献

相似文献

提出了一种在中文文档图像中搜索用户指定的词/短语的方法,该方法不需要进行版面分析。首先利用连通分量分析确定汉字图像的包围盒。接下来,从用户指定的单词/短语中选择合适的字符作为初始字符,以在文档中搜索匹配的候选。一旦找到匹配的候选,则在位置关系和大小相似性的约束下,检查其在水平和垂直方向上的相邻字符是否与用户指定的单词/短语中的其他对应字符匹配。字符匹配分两个阶段完成。基于笔画密度特征进行粗匹配。对于第二匹配阶段,提出了加权Hausdorff距离。实验结果表明,该方法能够有效地从文档图像的水平或垂直文本行中搜索到用户指定的中文单词/短语。
An approach to searching user-specified words/phases in Chinese document images, without the requirements of layout analysis, is proposed in this paper. Bounding boxes of Chinese character images are first determined using the connected component analysis. Next, a suitable character from the user-specified word/phrase is chosen as the initial character to search for a matching candidate in the document. Once a matched candidate is found, its adjacent characters in the horizontal and vertical directions are examined for matching with other corresponding characters in the user-specified word/phrase, subject to the constraints of positional relation and size similarity. The character matching is done in two stages. The coarse matching is carried out based on the stroke density features. A weighted Hausdorff distance is proposed for the second matching phase. Experimental results show that the proposed method can effectively search the user-specified Chinese word/phrase from horizontal or vertical text lines of document images.