DeepBIBX: Deep Learning for Image Based Bibliographic Data Extraction

DeepBIBX: Deep Learning for Image Based Bibliographic Data Extraction
复制标题

DeepBIBX:基于图像的书目数据提取的深度学习

DOI:
10.1007/978-3-319-70096-0_30
复制
发表时间:
2017
期刊:
影响因子:
--
通讯作者:
Ahmed S.
Ahmed S.
中科院分区:
--
文献类型:
--
作者:
Bhardwaj A;Mercier D;Dengel A;Ahmed S.

文献摘要

被引文献

相似文献

从非原生数字学术内容的文献图像中提取结构化书目数据是图书馆编目系统自动化和参考文献链接领域中一个具有挑战性的问题。现有的方法放弃了视觉线索,专注于将文档图像转换为文本,并使用训练好的分割模型进一步识别引文字符串。除了现有的方法需要大量的训练数据外,它们还依赖于语言。本文提出了一种新颖的方法(DeepBIBX),它从计算机视觉的角度来解决这个问题,并使用深度学习对文档图像中的单个引文字符串进行语义分割。DeepBIBX基于深度全卷积网络,并使用迁移学习从文档图像中提取书目参考文献。与现有的使用文本内容对书目参考进行语义分割的方法不同,DeepBIBX利用基于图像的上下文信息,这使得它适用于任何语言的文档。为了衡量所提出的方法的性能,收集了包含286个文档图像的数据集,其中包含5090个书目参考文献。评估结果表明,DeepBIBX在书目参考文献提取方面优于最先进的方法(ParsCit, 71.7%),达到84.9%的准确率,而ParsCit的准确率为71.7%。此外,在像素分类任务方面,DeepBIBX的准确率和召回率分别达到96.2%和94.4%。
Extraction of structured bibliographic data from document images of non-native-digital academic content is a challenging problem that finds its application in the automation of cataloging systems in libraries and reference linking domain. The existing approaches discard the visual cues and focus on converting the document image to text and further identifying citation strings using trained segmentation models. Apart from the large training data, which these existing methods require, they are also language dependent. This paper presents a novel approach (DeepBIBX) which targets this problem from a computer vision perspective and uses deep learning to semantically segment the individual citation strings in a document image. DeepBIBX is based on deep Fully Convolutional Networks and uses transfer learning to extract bibliographic references from document images. Unlike existing approaches which use textual content to semantically segment bibliographic references, DeepBIBX utilizes image based contextual information, which makes it applicable to documents of any language. To gauge the performance of the presented approach, a dataset consisting of 286 document images containing 5090 bibliographic references is collected. Evaluation results reveals that the DeepBIBX outperforms state-of-the-art method (ParsCit, 71.7%) for bibliographic references extraction and achieved an accuracy of 84.9% in comparison to 71.7%. Furthermore, in terms of pixel classification task, DeepBIBX achieved a precision and a recall rate of 96.2%, 94.4% respectively.