SideNoter: Scholarly Paper Browsing System based on PDF Restructuring and Text Annotation

SideNoter: Scholarly Paper Browsing System based on PDF Restructuring and Text Annotation
复制标题

SideNoter:基于PDF重构和文本注释的学术论文浏览系统

DOI:
--
复制
发表时间:
2016
期刊:
International Conference on Computational Linguistics
影响因子:
--
通讯作者:
Akiko Aizawa
Akiko Aizawa
中科院分区:
--
文献类型:
--
作者:
Takeshi Abekawa;Akiko Aizawa

文献摘要

被引文献

相似文献

在本文中,我们讨论了我们正在努力构建一个科学论文浏览系统,帮助用户阅读和理解先进的技术内容分布在PDF格式。由于PDF是一种专门为打印设计的格式,因此文档的布局和逻辑结构不可分割地嵌入到文件中。从PDF文件中提取自然语言文本需要花费大量精力,并在原始页面布局上显示NLP工具生成的语义注释。在我们的浏览系统中,我们解决了这些问题所造成的差距之间的打印文档和纯文本。我们的系统提供了从PDF文件中提取自然语言句子及其逻辑结构的方法,并将任意文本跨度映射到页面图像上的相应区域。我们使用ACL选集上发表的论文建立了一个演示系统,并演示了增强的搜索和改进的推荐功能,我们计划将其广泛提供给NLP研究人员。
In this paper, we discuss our ongoing efforts to construct a scientific paper browsing system that helps users to read and understand advanced technical content distributed in PDF. Since PDF is a format specifically designed for printing, layout and logical structures of documents are indistinguishably embedded in the file. It requires much effort to extract natural language text from PDF files, and reversely, display semantic annotations produced by NLP tools on the original page layout. In our browsing system, we tackle these issues caused by the gap between printable document and plain text. Our system provides ways to extract natural language sentences from PDF files together with their logical structures, and also to map arbitrary textual spans to their corresponding regions on page images. We setup a demonstration system using papers published in ACL anthology and demonstrate the enhanced search and refined recommendation functions which we plan to make widely available to NLP researchers.