Layout-aware text extraction from full-text PDF of scientific articles.

Layout-aware text extraction from full-text PDF of scientific articles.
复制标题

DOI:
10.1186/1751-0473-7-7
复制
发表时间:
2012-05-28
影响因子:
--
通讯作者:
Burns GA
Burns GA
中科院分区:
其他
文献类型:
--
作者:
Ramakrishnan C;Patnia A;Hovy E;Burns GA

文献摘要

被引文献

相似文献

可移植文档格式(PDF)是在线科学出版物最常用的文件格式。缺乏有效的手段来提取文本从这些PDF文件中的布局感知的方式提出了一个重大的挑战,生物医学文本挖掘或biocuration信息系统的开发人员使用出版的文献作为信息源。在本文中,我们介绍了“布局感知PDF文本提取”(LA-PDFText)系统,以方便准确提取文本从PDF文件的研究文章中使用的文本挖掘应用程序。我们的论文介绍了一个开源系统的建设和性能,提取PDF格式的全文研究文章的文本块,并将其分类为逻辑单元的基础上,具有特定部分的规则。LA-PDFText系统只关注研究文章的文本内容,并作为进一步实验的基线,以处理多模态内容(如图像和图形)的更高级提取方法。该系统工作在一个三阶段的过程中:(1)检测连续的文本块使用空间布局处理定位和识别连续的文本块,(2)分类文本块到修辞类别使用基于规则的方法和(3)拼接分类文本块在一起,在正确的顺序,导致从部分的文本块提取。我们表明,我们的系统可以识别文本块,并将其分为修辞类与精确度1 = 0.96%召回= 0.89%和F1 = 0.91%。我们还提出了在步骤2中使用的块检测算法的准确性的评估。此外,我们还比较了LA-PDFText提取的文本与PubMed Central开放获取子集的文本的准确性。然后,我们将此准确性与PDF 2 Text系统提取的文本的准确性进行了比较,该系统2通常用于从PDF中提取文本。最后,我们讨论了初步的错误分析,我们的系统,并确定进一步改进的领域。LA-PDFText是一个开源工具,用于从全文科学文章中准确提取文本。该系统的发布可在http://code.google.com/p/lapdftext/上获得。
The Portable Document Format (PDF) is the most commonly used file format for online scientific publications. The absence of effective means to extract text from these PDF files in a layout-aware manner presents a significant challenge for developers of biomedical text mining or biocuration informatics systems that use published literature as an information source. In this paper we introduce the ‘Layout-Aware PDF Text Extraction’ (LA-PDFText) system to facilitate accurate extraction of text from PDF files of research articles for use in text mining applications. Our paper describes the construction and performance of an open source system that extracts text blocks from PDF-formatted full-text research articles and classifies them into logical units based on rules that characterize specific sections. The LA-PDFText system focuses only on the textual content of the research articles and is meant as a baseline for further experiments into more advanced extraction methods that handle multi-modal content, such as images and graphs. The system works in a three-stage process: (1) Detecting contiguous text blocks using spatial layout processing to locate and identify blocks of contiguous text, (2) Classifying text blocks into rhetorical categories using a rule-based method and (3) Stitching classified text blocks together in the correct order resulting in the extraction of text from section-wise grouped blocks. We show that our system can identify text blocks and classify them into rhetorical categories with Precision1 = 0.96% Recall = 0.89% and F1 = 0.91%. We also present an evaluation of the accuracy of the block detection algorithm used in step 2. Additionally, we have compared the accuracy of the text extracted by LA-PDFText to the text from the Open Access subset of PubMed Central. We then compared this accuracy with that of the text extracted by the PDF2Text system, 2commonly used to extract text from PDF. Finally, we discuss preliminary error analysis for our system and identify further areas of improvement. LA-PDFText is an open-source tool for accurately extracting text from full-text scientific articles. The release of the system is available at http://code.google.com/p/lapdftext/.