Error Detection of CRF-Based Bibliography Extraction from Reference Strings

Error Detection of CRF-Based Bibliography Extraction from Reference Strings
复制标题

DOI:
10.1007/978-3-642-34752-8_29
复制
发表时间:
2012-11
期刊:
--
影响因子:
--
通讯作者:
Manabu Ohta;Daiki Arauchi;A. Takasu;J. Adachi
Manabu Ohta;Daiki Arauchi;A. Takasu;J. Adachi
中科院分区:
其他
文献类型:
--
作者:
Manabu Ohta;Daiki Arauchi;A. Takasu;J. Adachi

文献摘要

相似文献

我们提出了一种对通常列在研究论文末尾的参考文献字符串进行解析的方法,以从中提取重要的参考书目,如标题。该方法使用条件随机场(CRF)来估计从参考字符串生成的令牌序列中每个令牌的正确书目标签。虽然我们对一份日本学术期刊的句法分析达到了合理的精度,但错误是不可避免的。因此,本文提出了提高基于CRF的书目解析的可信度的方法,以检测此类解析错误。本文还报告了对所提出的句法分析的经验评估,不仅基于其准确性,而且还基于其检测错误的难易程度。实验表明,所提出的方法合理地指出了句法错误,可以在适度的人工后期编辑代价下提高摘录书目的质量。
We proposed a parsing method for reference strings usually listed at the end of research papers to extract important bibliographies such as a title from them. The method uses a conditional random field (CRF) to estimate the correct bibliographic label for each token in the token sequence generated from a reference string. Although we achieved reasonable parsing accuracies for a Japanese academic journal, errors are inevitable. Therefore, this paper proposes ways to increase confidence for CRF-based bibliography parsing to detect such parsing errors. This paper also reports an empirical evaluation of the proposed parsing on the basis not only of its accuracies but also of how easy it is to detect errors. The experiments showed that the proposed measures reasonably indicated parsing errors and could be used to improve the quality of extracted bibliographies at a moderate manual post-editing cost.