Segmentation of Heterogeneous Documents into Homogeneous Components using Morphological Operations

Segmentation of Heterogeneous Documents into Homogeneous Components using Morphological Operations
复制标题

使用形态学操作将异构文档分割为同质组件

DOI:
10.1109/icis.2018.8466395
复制
发表时间:
2018
期刊:
2018 IEEE/ACIS 17th International Conference on Computer and Information Science (ICIS)
影响因子:
--
通讯作者:
Hasnain Heickal
Hasnain Heickal
中科院分区:
--
文献类型:
--
作者:
Nasid Habib Barna;Tisa Islam Erana;Shabbir Ahmed;Hasnain Heickal

文献摘要

被引文献

相似文献

近年来,文档版面分析的研究在很大程度上得到了广泛的应用,人们对文档版面分析效率的要求也日益提高。在分析版面之前,文档分割是一个重要的预处理步骤。本文提出了一种独立于语言的文档分割系统,该系统将一个不同种类的打印文档分割成同质部分,如半色调和图形、文本和表格及其单个单元。从一个输入文档页面中,同质部分被分成三个步骤,分别是半色调图像的提取、表格的提取和文本块的分割。这些模块共同构建了整个页面分割系统,该系统获取异质文档页面的输入图像,并生成带有彩色边界框的明确指示的同质片段的输出。这些模块使用形态运算来检测组件。为了提高图像分割的性能,提出了残差图像片段检索算法(RIFR)。本文还提出了从表格单元格中提取文本的方法。将RIFR和TETC结合起来,我们得到了93%的总体准确率。表格和单元格检测的准确率较高,达到96%,而图像和文本检测的准确率约为90%。
The research on document layout analysis has been widespread over a large arena recently and is craving for more efficiency day by day. Document segmentation is an important preprocessing step before analyzing the layouts. This paper presents a language-independent document segmentation system that segments a heterogeneous printed document into homogeneous components like halftones and graphics, texts and tables including its individual cells. From an input document page homogeneous components are segmented in three steps with three separate modules, which are- extraction of halftone images, extraction of tables and segmentation of text blocks. These modules altogether build the whole page segmentation system which takes an input image of heterogeneous document page and produces an output with explicitly indicated homogeneous segments with colored bounding boxes. The modules use morphological operations to detect the components. To improve the performance of image segmentation Residual Image Fragments Retrieval (RIFR) is proposed. The paper also proposes Text Extraction from Table Cells (TETC). Combining RIFR and TETC together we get an overall accuracy of 93%. Table and cell detection have a higher accuracy of 96% whereas image and texts have around 90% accuracy.