Improving Information Extraction from Visually Rich Documents using Visual Span Representations

Improving Information Extraction from Visually Rich Documents using Visual Span Representations
复制标题

DOI:
10.14778/3446095.3446104
复制
发表时间:
2021-01
期刊:
Proc. VLDB Endow.
影响因子:
--
通讯作者:
Ritesh Sarkhel;Arnab Nandi
Ritesh Sarkhel;Arnab Nandi
中科院分区:
其他
文献类型:
--
作者:
Ritesh Sarkhel;Arnab Nandi

文献摘要

相似文献

与文本内容一样,视觉特征在视觉丰富的文档的语义中发挥着重要作用。如果不考虑这些视觉提示,信息提取 (IE) 任务在这些文档上的表现就会很差。在本文中,我们提出了 Artemis——一种视觉感知、基于机器学习的 IE 方法,用于处理异构视觉丰富的文档。 Artemis 通过联合编码 IE 任务的视觉和文本上下文来表示文档中的视觉跨度。我们的主要贡献有两个。首先,我们开发了一个深度学习模型,可以用最少的人工标签来识别视觉范围的局部上下文边界。其次,我们描述了一种深度神经网络,该网络通过考虑文本和布局特定的特征,将视觉跨度的多模态上下文编码为固定长度的向量。它通过利用学习到的表示和推理任务来识别包含命名实体的视觉范围。我们通过一系列信息提取任务在来自不同领域的四个异构数据集上评估 Artemis。结果表明,它的 F1 分数比最先进的基于文本的方法高出 17 分。
Along with textual content, visual features play an essential role in the semantics of visually rich documents. Information extraction (IE) tasks perform poorly on these documents if these visual cues are not taken into account. In this paper, we present Artemis - a visually aware, machine-learning-based IE method for heterogeneous visually rich documents. Artemis represents a visual span in a document by jointly encoding its visual and textual context for IE tasks. Our main contribution is two-fold. First, we develop a deep-learning model that identifies the local context boundary of a visual span with minimal human-labeling. Second, we describe a deep neural network that encodes the multimodal context of a visual span into a fixed-length vector by taking its textual and layout-specific features into account. It identifies the visual span(s) containing a named entity by leveraging this learned representation followed by an inference task. We evaluate Artemis on four heterogeneous datasets from different domains over a suite of information extraction tasks. Results show that it outperforms state-of-the-art text-based methods by up to 17 points in F1-score.