Scholarly big data quality assessment: a case study of document linking and conflation with S2ORC

Scholarly big data quality assessment: a case study of document linking and conflation with S2ORC
复制标题

DOI:
10.1145/3558100.3563850
复制
发表时间:
2022-09
期刊:
Proceedings of the 22nd ACM Symposium on Document Engineering
影响因子:
--
通讯作者:
Jian Wu;Ryan Hiltabrand;Dominik Soós;C. Lee Giles
Jian Wu;Ryan Hiltabrand;Dominik Soós;C. Lee Giles
中科院分区:
其他
文献类型:
--
作者:
Jian Wu;Ryan Hiltabrand;Dominik Soós;C. Lee Giles

文献摘要

相似文献

最近,艾伦人工智能研究所发布了语义学者开放研究语料库(S2 ORC),这是最大的开放获取学术大数据集之一,拥有超过1.3亿条学术论文记录。S2ORC包含自动生成的元数据的重要部分。元数据的质量可能会影响下游任务,如引文分析、引文预测和链接分析。在这个项目中,我们评估的文件链接质量和估计的文件合并率为S2ORC数据集。使用半自动策划的地面实况语料库,我们估计整体文档链接质量很高,92.6%的文档正确链接到六个主要数据库,但链接质量因主题领域而异。文档合并率约为2.6%,这意味着约97.4%的文档是唯一的。我们进一步使用从S2ORC创建的地面实况定量比较了三种近似重复检测方法。实验表明,位置敏感哈希是最好的方法,在有效性和可扩展性,实现高性能(F1=0.960)和大大减少运行时间。我们的代码和数据可在https://github.com/lamps-lab/docconflation上获得。
Recently, the Allen Institute for Artificial Intelligence released the Semantic Scholar Open Research Corpus (S2ORC), one of the largest open-access scholarly big datasets with more than 130 million scholarly paper records. S2ORC contains a significant portion of automatically generated metadata. The metadata quality could impact downstream tasks such as citation analysis, citation prediction, and link analysis. In this project, we assess the document linking quality and estimate the document conflation rate for the S2ORC dataset. Using semi-automatically curated ground truth corpora, we estimated that the overall document linking quality is high, with 92.6% of documents correctly linking to six major databases, but the linking quality varies depending on subject domains. The document conflation rate is around 2.6%, meaning that about 97.4% of documents are unique. We further quantitatively compared three near-duplicate detection methods using the ground truth created from S2ORC. The experiments indicated that locality-sensitive hashing was the best method in terms of effectiveness and scalability, achieving high performance (F1=0.960) and a much reduced runtime. Our code and data are available at https://github.com/lamps-lab/docconflation.