BROS: A Pre-trained Language Model Focusing on Text and Layout for Better Key Information Extraction from Documents

BROS: A Pre-trained Language Model Focusing on Text and Layout for Better Key Information Extraction from Documents
复制标题

BROS:一种预训练的语言模型,专注于文本和布局,以便更好地从文档中提取关键信息

DOI:
10.1609/aaai.v36i10.21322
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
Sungrae Park
Sungrae Park
中科院分区:
--
文献类型:
--
作者:
Teakgyu Hong;Donghyun Kim;Mingi Ji;Wonseok Hwang;Daehyun Nam;Sungrae Park

文献摘要

参考文献

被引文献

相似文献

从文档图像中提取关键信息(KIE)需要理解文本在二维空间中的上下文和空间语义。 最近的许多研究试图通过开发预先训练的语言模型来解决这一问题,该模型侧重于将文档图像的视觉特征与文本及其布局相结合。 另一方面,本文从文本与版面的有效结合入手,从根本上来解决这一问题。 具体地说,我们提出了一种预先训练的语言模型Bros(BERT Depending On Spatiality),该模型对文本在2D空间中的相对位置进行编码,并使用区域掩蔽策略从未标记的文档中学习。 通过这种优化的2D空间文本理解训练方案,Bros在不依赖视觉特征的情况下,在四个KIE基准(FUNSD、SROIE*、CORD和SciTSR)上表现出与以前的方法相当或更好的性能。 本文还揭示了KIE任务中的两个现实挑战--(1)最小化错误文本排序的错误和(2)从较少的下游示例中进行有效学习--并证明了Bros方法比以往方法的优越性。
Key information extraction (KIE) from document images requires understanding the contextual and spatial semantics of texts in two-dimensional (2D) space. Many recent studies try to solve the task by developing pre-trained language models focusing on combining visual features from document images with texts and their layout. On the other hand, this paper tackles the problem by going back to the basic: effective combination of text and layout. Specifically, we propose a pre-trained language model, named BROS (BERT Relying On Spatiality), that encodes relative positions of texts in 2D space and learns from unlabeled documents with area-masking strategy. With this optimized training scheme for understanding texts in 2D space, BROS shows comparable or better performance compared to previous methods on four KIE benchmarks (FUNSD, SROIE*, CORD, and SciTSR) without relying on visual features. This paper also reveals two real-world challenges in KIE tasks--(1) minimizing the error from incorrect text ordering and (2) efficient learning from fewer downstream examples--and demonstrates the superiority of BROS over previous methods.
DOI: 10.1007/978-3-030-86331-9_47
发表时间: 2021-02
期刊: ArXiv
影响因子: --
作者:
Rafal Powalski;Łukasz Borchmann;Dawid Jurkiewicz;Tomasz Dwojak;Michal Pietruszka;Gabriela Pałka
通讯作者: Rafal Powalski;Łukasz Borchmann;Dawid Jurkiewicz;Tomasz Dwojak;Michal Pietruszka;Gabriela Pałka