Automatic identification of academic articles in Japanese PDF files
Automatic identification of academic articles in Japanese PDF files
复制标题
自动识别日语PDF文件中的学术文章
DOI:
10.46895/lis.56.43
复制
发表时间:
2006
影响因子:
0.5
通讯作者:
S. Ueda
中科院分区:
文献类型:
--
作者:
Teru Agata;Atsushi Ikeuchi;Emi Ishita;Michiko Nozue;T. Kuno;S. Ueda
As open-access becomes common, many researchers deposit their research products in a publicly accessible web (i.e. self-archiving). Although they are accessible from general search engines, massive other contents tend to hide them. The purpose of this research is to identify academic articles or quasi-articles from the entire web automatically. In this paper we conduct experiments on the performance of various classifiers and compare in terms of precision, recall, F-value. The classifiers used such attributes as terms appeared in PDF files and empirical rules. The diverse performance of each classifier discloses its characteristics.