Representation Learning for Information Extraction from Form-like Documents

Representation Learning for Information Extraction from Form-like Documents
复制标题

从类似表单的文档中提取信息的表示学习

DOI:
10.18653/v1/2020.acl-main.580
复制
发表时间:
2020
期刊:
2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Marc Najork
Marc Najork
中科院分区:
--
文献类型:
--
作者:
Bodhisattwa Prasad Majumder;Navneet Potti;Sandeep Tata;James Bradley Wendt;Qi Zhao;Marc Najork

文献摘要

参考文献

被引文献

相似文献

我们提出了一种新的基于表征学习的方法来解决从表格类文档图像中提取结构化信息的问题。我们提出了一个使用目标领域类型的知识来生成抽取候选者的抽取系统,以及一个基于文档中相邻单词学习每个候选者的密集表示的神经网络结构。这些学习的表示不仅有助于解决来自两个不同域的不可见文档模板的提取任务,而且也是可解释的,如我们使用丢失案例所示。
We propose a novel approach using representation learning for tackling the problem of extracting structured information from form-like document images. We propose an extraction system that uses knowledge of the types of the target fields to generate extraction candidates and a neural network architecture that learns a dense representation of each candidate based on neighboring words in the document. These learned representations are not only useful in solving the extraction task for unseen document templates from two different domains but are also interpretable, as we show using loss cases.
DOI: 10.1145/775152.775155
发表时间: 2003-05
期刊: --
影响因子: --
作者:
Shipeng Yu;Deng Cai;Ji-Rong Wen;Wei-Ying Ma
通讯作者: Shipeng Yu;Deng Cai;Ji-Rong Wen;Wei-Ying Ma