Identifying References to Datasets in Publications

Identifying References to Datasets in Publications
复制标题

识别出版物中对数据集的引用

DOI:
10.1007/978-3-642-33290-6_17
复制
发表时间:
2012
期刊:
Science, Technology, & Human Values
影响因子:
--
通讯作者:
Brigitte Mathiak
Brigitte Mathiak
中科院分区:
--
文献类型:
--
作者:
K. Boland;Dominique Ritze;K. Eckert;Brigitte Mathiak

文献摘要

被引文献

相似文献

研究数据和出版物通常存储在单独的和结构不同的信息系统中。通常,这些资源之间的联系并不明确,这使得以前的研究变得复杂。本文提出了一种基于模式归纳的全文参考文献检测方法。由于这些引用没有以标准化的方式指定,并且可能出现在各种不同的上下文中,即,标题,脚注,或连续文本-我们的算法需要诱导非常灵活的模式。为了克服训练样本的稀疏分布,我们使用自举方法迭代地诱导模式。我们表明,我们的方法取得了可喜的成果,自动识别的数据参考,是建立一个综合信息系统的第一步。
Research data and publications are usually stored in separate and structurally distinct information systems. Often, links between these resources are not explicitly available which complicates the search for previous research. In this paper, we propose a pattern induction method for the detection of study references in full texts. Since these references are not specified in a standardized way and may occur inside a variety of different contexts --- i.e., captions, footnotes, or continuous text --- our algorithm is required to induce very flexible patterns. To overcome the sparse distribution of training instances, we induce patterns iteratively using a bootstrapping approach. We show that our method achieves promising results for the automatic identification of data references and is a first step towards building an integrated information system.