Improving the accuracy of co-citation clustering using full text

Improving the accuracy of co-citation clustering using full text
复制标题

DOI:
10.1002/asi.22896
复制
发表时间:
2013-09-01
影响因子:
--
通讯作者:
Klavans, Richard
Klavans, Richard
中科院分区:
其他
文献类型:
--
作者:
Boyack, Kevin W.;Small, Henry;Klavans, Richard

文献摘要

被引文献

相似文献

历史上,共引模型仅基于书目信息。全文分析提供了显著提高这些共引模型所基于的信号质量的机会。在这项工作中,我们研究的影响,参考文献的接近度的准确性,共引集群。使用2007年的270,521篇全文文档的语料库,我们将仅使用书目信息的传统共引聚类的结果与共引聚类的结果进行比较,其中参考对之间的接近度被考虑到成对关系中。我们发现,占参考接近全文可以增加文本的连贯性(准确性的衡量标准)的共引聚类解决方案的30%,比传统的方法的基础上书目信息。
Historically, co-citation models have been based only on bibliographic information. Full-text analysis offers the opportunity to significantly improve the quality of the signals upon which these co-citation models are based. In this work we study the effect of reference proximity on the accuracy of co-citation clusters. Using a corpus of 270,521 full text documents from 2007, we compare the results of traditional co-citation clustering using only the bibliographic information to results from co-citation clustering where proximity between reference pairs is factored into the pairwise relationships. We find that accounting for reference proximity from full text can increase the textual coherence (a measure of accuracy) of a co-citation cluster solution by up to 30% over the traditional approach based on bibliographic information.