Agreement, the F-measure, and reliability in information retrieval
Agreement, the F-measure, and reliability in information retrieval
复制标题
DOI:
10.1197/jamia.m1733
复制
发表时间:
2005-05-01
影响因子:
6.4
通讯作者:
Rothschild, AS
中科院分区:
文献类型:
--
作者:
Hripcsak, G;Rothschild, AS
Information retrieval studies that involve searching the Internet or marking phrases usually lack a well-defined number of negative cases. This prevents the use of traditional interrater reliability metrics like the K statistic to assess the quality of expert-generated gold standards. Such studies often quantify system performance as precision, recall, and F-measure, or as agreement. It can be shown that the average F-measure among pairs of experts is numerically identical to the average positive specific agreement among experts and that K approaches these measures as the number of negative cases grows large. Positive specific agreement-or the equivalent F-measure may be an appropriate way to quantify interrater reliability and therefore to assess the reliability of a gold standard in these studies.