Automatic identification of academic articles in Japanese PDF files

Automatic identification of academic articles in Japanese PDF files
复制标题

自动识别日语PDF文件中的学术文章

DOI:
10.46895/lis.56.43
复制
发表时间:
2006
影响因子:
0.5
通讯作者:
S. Ueda
S. Ueda
中科院分区:
管理学4区
文献类型:
--
作者:
Teru Agata;Atsushi Ikeuchi;Emi Ishita;Michiko Nozue;T. Kuno;S. Ueda

文献摘要

被引文献

相似文献

随着开放获取变得普遍,许多研究人员将他们的研究成果存放在一个公共可访问的网络上(即自我存档)。尽管它们可以通过通用搜索引擎访问,但海量的其他内容往往会隐藏它们。本研究的目的是从整个网络中自动识别学术文章或准文章。本文对各种分类器的性能进行了实验,并从准确率、召回率、F值等方面进行了比较。分类器使用了出现在PDF文件中的术语和经验规则等属性。每个分类器的不同性能揭示了它的特点。
As open-access becomes common, many researchers deposit their research products in a publicly accessible web (i.e. self-archiving). Although they are accessible from general search engines, massive other contents tend to hide them. The purpose of this research is to identify academic articles or quasi-articles from the entire web automatically. In this paper we conduct experiments on the performance of various classifiers and compare in terms of precision, recall, F-value. The classifiers used such attributes as terms appeared in PDF files and empirical rules. The diverse performance of each classifier discloses its characteristics.