The gene normalization task in BioCreative III.

The gene normalization task in BioCreative III.
复制标题

DOI:
10.1186/1471-2105-12-s8-s2
复制
发表时间:
2011-10-03
期刊:
影响因子:
3
通讯作者:
Wilbur WJ
Wilbur WJ
中科院分区:
生物学4区
文献类型:
--
作者:
Lu Z;Kao HY;Wei CH;Huang M;Liu J;Kuo CJ;Hsu CN;Tsai RT;Dai HJ;Okazaki N;Cho HC;Gerner M;Solt I;Agarwal S;Liu F;Vishnyakova D;Ruch P;Romacker M;Rinaldi F;Bhattacharya S;Srinivasan P;Liu H;Torii M;Matos S;Campos D;Verspoor K;Livingston KM;Wilbur WJ

文献摘要

被引文献

相似文献

我们在 BioCreative III 中报告了基因标准化 (GN) 挑战,其中要求参赛团队返回全文文章中检测到的基因标识符的排名列表。为了进行培训,准备了 32 篇完整注释的文章和 500 篇部分注释的文章。总共选择了 507 篇文章作为测试集。由于注释成本高昂,不可能为所有测试文章获得黄金标准的人工注释。相反,我们开发了一种期望最大化(EM)算法方法,用于选择少量最有能力区分团队绩效的测试文章进行手动注释。此外,相同的算法随后被用于仅根据团队提交的内容推断基本事实。我们使用新提出的称为阈值平均精度 (TAP-k) 的指标来报告团队在黄金标准和推断的基本事实方面的表现。我们针对该任务总共收到了来自 14 个不同团队的 37 次运行。当使用 50 篇文章的黄金标准注释进行评估时,最高 TAP-k 分数分别为 0.3297 (k=5)、0.3538 (k=10) 和 0.3535 (k=20)。当在整个测试集上使用推断的基本事实进行评估时,观察到更高的 TAP-k 分数 0.4916(k=5、10、20)。当使用机器学习结合团队结果时,最佳复合系统在黄金标准上获得了 0.3707 (k=5)、0.4311 (k=10) 和 0.4477 (k=20) 的 TAP-k 分数,分别比最佳团队结果提高了 12.4%、21.8% 和 26.6%。通过使用全文和非特定物种,BioCreative III 中的 GN 任务比过去的类似任务更接近真正的文献管理任务,并为文本挖掘社区带来了额外的挑战,正如整体团队结果所揭示的那样。通过使用黄金标准评估团队,我们表明 EM 算法允许区分团队提交的内容,同时保持手动注释工作的可行性。使用推断的基本事实,我们展示了团队之间比较绩效的衡量标准。最后,通过比较黄金标准与推断的基本事实的团队排名,我们进一步证明推断的基本事实与检测良好团队绩效的黄金标准一样有效。
We report the Gene Normalization (GN) challenge in BioCreative III where participating teams were asked to return a ranked list of identifiers of the genes detected in full-text articles. For training, 32 fully and 500 partially annotated articles were prepared. A total of 507 articles were selected as the test set. Due to the high annotation cost, it was not feasible to obtain gold-standard human annotations for all test articles. Instead, we developed an Expectation Maximization (EM) algorithm approach for choosing a small number of test articles for manual annotation that were most capable of differentiating team performance. Moreover, the same algorithm was subsequently used for inferring ground truth based solely on team submissions. We report team performance on both gold standard and inferred ground truth using a newly proposed metric called Threshold Average Precision (TAP-k). We received a total of 37 runs from 14 different teams for the task. When evaluated using the gold-standard annotations of the 50 articles, the highest TAP-k scores were 0.3297 (k=5), 0.3538 (k=10), and 0.3535 (k=20), respectively. Higher TAP-k scores of 0.4916 (k=5, 10, 20) were observed when evaluated using the inferred ground truth over the full test set. When combining team results using machine learning, the best composite system achieved TAP-k scores of 0.3707 (k=5), 0.4311 (k=10), and 0.4477 (k=20) on the gold standard, representing improvements of 12.4%, 21.8%, and 26.6% over the best team results, respectively. By using full text and being species non-specific, the GN task in BioCreative III has moved closer to a real literature curation task than similar tasks in the past and presents additional challenges for the text mining community, as revealed in the overall team results. By evaluating teams using the gold standard, we show that the EM algorithm allows team submissions to be differentiated while keeping the manual annotation effort feasible. Using the inferred ground truth we show measures of comparative performance between teams. Finally, by comparing team rankings on gold standard vs. inferred ground truth, we further demonstrate that the inferred ground truth is as effective as the gold standard for detecting good team performance.