Assessor disagreement and text classifier accuracy

Assessor disagreement and text classifier accuracy
复制标题

评估者分歧和文本分类器准确性

DOI:
10.1145/2484028.2484156
复制
发表时间:
2013
期刊:
Proceedings of the 36th international ACM SIGIR conference on Research and development in information retrieval
影响因子:
--
通讯作者:
Jeremy Pickens
Jeremy Pickens
中科院分区:
--
文献类型:
--
作者:
William Webber;Jeremy Pickens

文献摘要

被引文献

相似文献

Text classifiers are frequently used for high-yield retrieval from large corpora, such as in e-discovery. The classifier is trained by annotating example documents for relevance. These examples may, however, be assessed by people other than those whose conception of relevance is authoritative. In this paper, we examine the impact that disagreement between actual and authoritative assessor has upon classifier effectiveness, when evaluated against the authoritative conception. We find that using alternative assessors leads to a significant decrease in binary classification quality, though less so ranking quality. A ranking consumer would have to go on average 25% deeper in the ranking produced by alternative-assessor training to achieve the same yield as for authoritative-assessor training.