Computational linguistics literature and citations oriented citation linkage, classification and summarization

Computational linguistics literature and citations oriented citation linkage, classification and summarization
复制标题

DOI:
10.1007/s00799-017-0219-5
复制
发表时间:
2018-09
影响因子:
1.5
通讯作者:
Lei Li;Liyuan Mao;Yazhao Zhang;Junqi Chi;Taiwen Huang;Xiaoyue Cong;Heng Peng
Lei Li;Liyuan Mao;Yazhao Zhang;Junqi Chi;Taiwen Huang;Xiaoyue Cong;Heng Peng
中科院分区:
--
文献类型:
--
作者:
Lei Li;Liyuan Mao;Yazhao Zhang;Junqi Chi;Taiwen Huang;Xiaoyue Cong;Heng Peng

文献摘要

相似文献

科学文献是当今学者最重要的资源,其引文为研究者分析科学发展趋势、影响以及作品与作者之间的关系提供了一个潜在的强有力的途径。本文主要研究计算语言学科学文献的自动引文分析和摘要,这也是2016年BIRNDL第二届计算语言学科学文献摘要研讨会(The Joint Workshop on Bibliometric-enhanced Information Retrieval and Natural Language Processing for Digital Libraries)的共同任务。通过各种计算方法,根据引文与参考文献中文本跨度之间的内容相似性识别引文与参考文献中文本跨度之间的每个引文链接。然后将引用文本跨度分类为五个预定义的方面,即,基于支持向量机和投票法的词汇和规则的各种特征,提出了假设、含义、目标、结果和方法。最后,从引用的文本跨度的参考文献的摘要是在250字内生成的。采用hLDA(hierarchical Latent Dirichlet Allocation)主题模型进行内容建模,为摘要提供句子聚类(子主题)和词分布(抽象性)知识。我们结合联合收割机hLDA知识与其他几个经典的功能,使用不同的权重和比例来评估参考文献中的句子。根据BIRNDL 2016发布的评估结果,我们的系统排名前一和前二,这验证了我们的方法的有效性。
Scientific literature is currently the most important resource for scholars, and their citations have provided researchers with a powerful latent way to analyze scientific trends, influences and relationships of works and authors. This paper is focused on automatic citation analysis and summarization for the scientific literature of computational linguistics, which are also the shared tasks in the 2016 workshop of the 2nd Computational Linguistics Scientific Document Summarization at BIRNDL 2016 (The Joint Workshop on Bibliometric-enhanced Information Retrieval and Natural Language Processing for Digital Libraries). Each citation linkage between a citation and the spans of text in the reference paper is recognized according to their content similarities via various computational methods. Then the cited text span is classified to five pre-defined facets, i.e., Hypothesis, Implication, Aim, Results and Method, based on various features of lexicons and rules via Support Vector Machine and Voting Method. Finally, a summary of the reference paper from the cited text spans is generated within 250 words. hLDA (hierarchical Latent Dirichlet Allocation) topic model is adopted for content modeling, which provides knowledge about sentence clustering (subtopic) and word distributions (abstractiveness) for summarization. We combine hLDA knowledge with several other classical features using different weights and proportions to evaluate the sentences in the reference paper. Our systems have been ranked top one and top two according to the evaluation results published by BIRNDL 2016, which has verified the effectiveness of our methods.