Visual Segmentation for Information Extraction from Heterogeneous Visually Rich Documents

Visual Segmentation for Information Extraction from Heterogeneous Visually Rich Documents
复制标题

DOI:
10.1145/3299869.3319867
复制
发表时间:
2019-06
期刊:
Proceedings of the 2019 International Conference on Management of Data
影响因子:
--
通讯作者:
Ritesh Sarkhel;Arnab Nandi
Ritesh Sarkhel;Arnab Nandi
中科院分区:
其他
文献类型:
--
作者:
Ritesh Sarkhel;Arnab Nandi

文献摘要

被引文献

相似文献

物理和数字文档通常包含视觉上丰富的信息。有了这些信息,文档中就没有严格的数据值必须出现的顺序或位置。除了文本线索外,这些文档通常还依赖于显著的视觉特征来定义不同的语义边界并增强它们传播的信息。在执行信息提取(IE)时,传统技术存在不足,因为它们只使用文本表示,而不考虑这些文档布局固有的视觉线索。我们提出了一种从异构视觉丰富文档中提取信息的通用方法VS2。这项工作有两个主要贡献。首先,我们提出了一种鲁棒分割算法,该算法将视觉丰富的文档分解为视觉隔离但语义连贯的区域,称为逻辑块。在这个过程中使用了与文档类型无关的低级视觉和语义特征。我们的第二个贡献是远程监督的搜索和选择方法,该方法利用这些逻辑块定义的上下文边界来标识这些文档中的命名实体。在三个异构数据集上的实验结果表明,该方法在所有数据集上的性能都明显优于纯文本方法。将其与最先进的方法进行比较还表明,VS2在所有数据集上的性能相当或更好。
Physical and digital documents often contain visually rich information. With such information, there is no strict ordering or positioning in the document where the data values must appear. Along with textual cues, these documents often also rely on salient visual features to define distinct semantic boundaries and augment the information they disseminate. When performing information extraction (IE), traditional techniques fall short, as they use a text-only representation and do not consider the visual cues inherent to the layout of these documents. We propose VS2, a generalized approach for information extraction from heterogeneous visually rich documents. There are two major contributions of this work. First, we propose a robust segmentation algorithm that decomposes a visually rich document into a bag of visually isolated but semantically coherent areas, called logical blocks. Document type agnostic low-level visual and semantic features are used in this process. Our second contribution is a distantly supervised search-and-select method for identifying the named entities within these documents by utilizing the context boundaries defined by these logical blocks. Experimental results on three heterogeneous datasets suggest that the proposed approach significantly outperforms its text-only counterparts on all datasets. Comparing it against the state-of-the-art methods also reveal that VS2 performs comparably or better on all datasets.