Agreement, the F-measure, and reliability in information retrieval

Agreement, the F-measure, and reliability in information retrieval
复制标题

DOI:
10.1197/jamia.m1733
复制
发表时间:
2005-05-01
影响因子:
6.4
通讯作者:
Rothschild, AS
Rothschild, AS
中科院分区:
管理学2区
文献类型:
--
作者:
Hripcsak, G;Rothschild, AS

文献摘要

被引文献

相似文献

涉及互联网搜索或短语标记的信息检索研究通常缺乏明确定义的否定案例数量。这阻止了使用传统的互判员可靠性指标,如K统计量来评估专家生成的黄金标准的质量。此类研究通常将系统性能量化为精度、召回率和f度量,或者一致性。可以证明,专家对之间的平均f测度与专家之间的平均正特定一致性在数值上是相同的,并且K随着负面案例数量的增加而接近这些测度。积极的具体协议-或等效的f -测量可能是量化互解释器可靠性的适当方法,因此可以评估这些研究中金标准的可靠性。
Information retrieval studies that involve searching the Internet or marking phrases usually lack a well-defined number of negative cases. This prevents the use of traditional interrater reliability metrics like the K statistic to assess the quality of expert-generated gold standards. Such studies often quantify system performance as precision, recall, and F-measure, or as agreement. It can be shown that the average F-measure among pairs of experts is numerically identical to the average positive specific agreement among experts and that K approaches these measures as the number of negative cases grows large. Positive specific agreement-or the equivalent F-measure may be an appropriate way to quantify interrater reliability and therefore to assess the reliability of a gold standard in these studies.