Efficient Knowledge Graph Accuracy Evaluation

Efficient Knowledge Graph Accuracy Evaluation
复制标题

DOI:
10.14778/3342263.3342642
复制
发表时间:
2019-07-01
影响因子:
2.5
通讯作者:
Yang, Jun
Yang, Jun
中科院分区:
计算机科学2区
文献类型:
--
作者:
Gao, Junyang;Li, Xian;Yang, Jun

文献摘要

被引文献

相似文献

对大规模知识图(KG)的准确性进行估计通常需要人类对图中的样本进行注释。如何在保持人工标注成本低的同时获得统计上有意义的准确性评估估计是KG开发周期及其实际应用的关键问题。令人惊讶的是,这个具有挑战性的问题在以前的研究中基本上被忽视了。为了解决这个问题,本文提出了一个有效的抽样和评估框架,其目的是提供质量的准确性评估与强有力的统计保证,同时最大限度地减少人为努力。出于在实践中观察到的注释成本函数的属性,我们建议使用聚类抽样来降低整体成本。我们进一步应用加权和两阶段抽样以及分层更好的抽样设计。我们还扩展了我们的框架,使有效的增量评估不断变化的KG,引入两个解决方案的基础上分层抽样和加权变量的油藏抽样。在真实数据集上的大量实验证明了我们提出的解决方案的有效性和效率。与基准方法相比,我们的最佳解决方案可以在静态KG评估上降低高达60%的成本,在不断发展的KG评估上降低高达80%的成本,而不会损失评估质量。
Estimation of the accuracy of a large-scale knowledge graph (KG) often requires humans to annotate samples from the graph. How to obtain statistically meaningful estimates for accuracy evaluation while keeping human annotation costs low is a problem critical to the development cycle of a KG and its practical applications. Surprisingly, this challenging problem has largely been ignored in prior research. To address the problem, this paper proposes an efficient sampling and evaluation framework, which aims to provide quality accuracy evaluation with strong statistical guarantee while minimizing human efforts. Motivated by the properties of the annotation cost function observed in practice, we propose the use of cluster sampling to reduce the overall cost. We further apply weighted and two-stage sampling as well as stratification for better sampling designs. We also extend our framework to enable efficient incremental evaluation on evolving KG, introducing two solutions based on stratified sampling and a weighted variant of reservoir sampling. Extensive experiments on real-world datasets demonstrate the effectiveness and efficiency of our proposed solution. Compared to baseline approaches, our best solutions can provide up to 60% cost reduction on static KG evaluation and up to 80% cost reduction on evolving KG evaluation, without loss of evaluation quality.