Tasks, topics and relevance judging for the TREC Genomics Track: five years of experience evaluating biomedical text information retrieval systems

Tasks, topics and relevance judging for the TREC Genomics Track: five years of experience evaluating biomedical text information retrieval systems
复制标题

DOI:
10.1007/s10791-008-9072-x
复制
发表时间:
2009-02-01
期刊:
INFORMATION RETRIEVAL
影响因子:
--
通讯作者:
Hersh, William R.
Hersh, William R.
中科院分区:
其他
文献类型:
--
作者:
Roberts, Phoebe M.;Cohen, Aaron M.;Hersh, William R.

文献摘要

被引文献

相似文献

在专家生物学家评委团队的帮助下,TREC基因组学轨道已经生成了四个大型的“黄金标准”测试集,包括100多个独特的主题、两种特别的检索任务以及它们相应的相关性判断。多年来,日益复杂的任务要求制定评判工具和培训准则,以适应来自各种生物科学专业背景的兼职短期工作人员团队,并解决评估过程的一致性和可重复性问题。在影响试卷使用的因素方面吸取了重要的经验教训,包括题目设计、法官提供的注释、用于确定和培训法官的方法,以及提供中央主持人“元法官”。
With the help of a team of expert biologist judges, the TREC Genomics track has generated four large sets of "gold standard'' test collections, comprised of over a hundred unique topics, two kinds of ad hoc retrieval tasks, and their corresponding relevance judgments. Over the years of the track, increasingly complex tasks necessitated the creation of judging tools and training guidelines to accommodate teams of part-time short-term workers from a variety of specialized biological scientific backgrounds, and to address consistency and reproducibility of the assessment process. Important lessons were learned about factors that influenced the utility of the test collections including topic design, annotations provided by judges, methods used for identifying and training judges, and providing a central moderator "meta-judge''.