Navigating Imprecision in Relevance Assessments on the Road to Total Recall: Roger and Me

Navigating Imprecision in Relevance Assessments on the Road to Total Recall: Roger and Me
复制标题

在通向全面回忆的道路上应对相关性评估的不精确性:罗杰和我

DOI:
--
复制
发表时间:
2017
期刊:
Annual International ACM SIGIR Conference on Research and Development in Information Retrieval
影响因子:
--
通讯作者:
Maura R. Grossman
Maura R. Grossman
中科院分区:
--
文献类型:
--
作者:
G. Cormack;Maura R. Grossman

文献摘要

被引文献

相似文献

技术辅助审查(“TAR”)系统寻求实现“完全召回”;也就是说,尽可能接近100%召回率和100%准确率的理想,同时最大限度地减少人类审查工作。文献报道使用相关反馈的TAR方法可以达到比Voorhees报告的65%的召回率和65%的准确率高得多的“检索性能的实际上限”。因为这是人类彼此同意的水平”(相关性判断和检索有效性测量的变化,2000年)。这项工作认为,为了建立-以及评估-接近100%召回率和100%准确率的TAR系统,有必要对人类评估进行建模,而不是作为绝对的基础事实,而是作为被称为“相关性”的无定形属性的间接指标。模型的选择既影响系统有效性的评估,也影响相关反馈的模拟。模型提出,更好地适应现有的数据比可靠的地面实况模型。这些模型提出了提高TAR系统有效性的方法,以便混合人机系统可以提高人类审查的准确性和效率。通过使用两个数据集模拟TAR来测试这一假设:TREC 4 AdHoc集合,以及由401,960封电子邮件组成的数据集,这些电子邮件由一个人Roger以其高级国家记录档案管理员的官方身份手动审查和分类。使用TREC 4数据的结果表明,TAR比两个独立的NIST评估员的评估实现了更高的召回率和更高的准确率,罗杰在原始审查后两年多对电子邮件数据集进行了盲判定,表明他可以实现相同的召回率和更好的准确率,而审查的电子邮件远远少于401,960封,如果他使用TAR来代替详尽的人工审查。
Technology-assisted review ("TAR") systems seek to achieve "total recall"; that is, to approach, as nearly as possible, the ideal of 100% recall and 100% precision, while minimizing human review effort. The literature reports that TAR methods using relevance feedback can achieve considerably greater than the 65% recall and 65% precision reported by Voorhees as the "practical upper bound on retrieval performance... since that is the level at which humans agree with one another" (Variations in Relevance Judgments and the Measurement of Retrieval Effectiveness, 2000). This work argues that in order to build - as well as to, evaluate - TAR systems that approach 100% recall and 100% precision, it is necessary to model human assessment, not as absolute ground truth, but as an indirect indicator of the amorphous property known as "relevance." The choice of model impacts both the evaluation of system effectiveness, as well as the simulation of relevance feedback. Models are presented that better fit available data than the infallible ground-truth model. These models suggest ways to improve TAR-system effectiveness so that hybrid human-computer systems can improve on both the accuracy and efficiency of human review alone. This hypothesis is tested by simulating TAR using two datasets: the TREC 4 AdHoc collection, and a dataset consisting of 401,960 email messages that were manually reviewed and classified by a single individual, Roger, in his official capacity as Senior State Records Archivist. The results using the TREC 4 data show that TAR achieves higher recall and higher precision than the assessments by either of two independent NIST assessors, and blind adjudication of the email dataset, conducted by Roger, more than two years after his original review, shows that he could have achieved the same recall and better precision, while reviewing substantially fewer than 401,960 emails, had he employed TAR in place of exhaustive manual review.