Data preparation and interannotator agreement: BioCreAtIvE task 1B.

Data preparation and interannotator agreement: BioCreAtIvE task 1B.
复制标题

DOI:
10.1186/1471-2105-6-s1-s12
复制
发表时间:
2005
期刊:
影响因子:
3
通讯作者:
Hirschman L
Hirschman L
中科院分区:
生物学4区
文献类型:
--
作者:
Colosimo ME;Morgan AA;Yeh AS;Colombe JB;Hirschman L

文献摘要

被引文献

相似文献

我们准备并评估了培训和测试材料,以评估分子生物学中的文本挖掘方法。评估的目的是评估自动化系统从PubMed摘要中生成独特基因标识符列表的能力,用于三种模型生物飞行,小鼠和酵母。本文描述了培训和测试的答案键的准备和评估。这些由摘要中发现的归一化基因名称的列表组成,这些名称通过适应模型有机体数据库中的完整期刊文章的基因列表而产生。对于训练数据集,将基因列表自动修剪,以删除摘要中未找到的基因名称。对于测试数据集,通过提供指南的注释者手动注释进一步完善了它。解释评估结果的关键步骤是评估数据制备的质量。我们通过仔细评估通道人的协议和使用答案汇总结果来提高最终测试数据集的质量来做到这一点。 在一个小数据集上的通道间分析表明,我们的蝇和酵母基因列表很好(87%和91%的三向一致性),但鼠标基因列表列表有许多冲突(主要是遗漏),这导致了错误(69%的通道互合数协议)。通过比较和汇总参与者系统的答案,我们能够对测试数据进行额外的检查;这使我们能够找到其他错误,尤其是在鼠标中。这导致酵母变化1%并飞行“黄金标准”答案键,但鼠标答案键的变化为8%。 我们发现,明确的注释指南以及仔细的通道实验非常重要,以验证生成的基因列表。同样,单独的摘要是识别纸张中基因的糟糕资源,仅包含全文中提到的一部分基因(苍蝇为25%,小鼠为36%)。我们发现,模型生物数据库与同义词数量以及与策展标准有关的模型生物数据库之间存在内在差异。最后,我们发现答案汇总要快得多,并且使我们能够识别出比跨托子分析更多的冲突基因。
We prepared and evaluated training and test materials for an assessment of text mining methods in molecular biology. The goal of the assessment was to evaluate the ability of automated systems to generate a list of unique gene identifiers from PubMed abstracts for the three model organisms Fly, Mouse, and Yeast. This paper describes the preparation and evaluation of answer keys for training and testing. These consisted of lists of normalized gene names found in the abstracts, generated by adapting the gene list for the full journal articles found in the model organism databases. For the training dataset, the gene list was pruned automatically to remove gene names not found in the abstract; for the testing dataset, it was further refined by manual annotation by annotators provided with guidelines. A critical step in interpreting the results of an assessment is to evaluate the quality of the data preparation. We did this by careful assessment of interannotator agreement and the use of answer pooling of participant results to improve the quality of the final testing dataset. Interannotator analysis on a small dataset showed that our gene lists for Fly and Yeast were good (87% and 91% three-way agreement) but the Mouse gene list had many conflicts (mostly omissions), which resulted in errors (69% interannotator agreement). By comparing and pooling answers from the participant systems, we were able to add an additional check on the test data; this allowed us to find additional errors, especially in Mouse. This led to 1% change in the Yeast and Fly "gold standard" answer keys, but to an 8% change in the mouse answer key. We found that clear annotation guidelines are important, along with careful interannotator experiments, to validate the generated gene lists. Also, abstracts alone are a poor resource for identifying genes in paper, containing only a fraction of genes mentioned in the full text (25% for Fly, 36% for Mouse). We found that there are intrinsic differences between the model organism databases related to the number of synonymous terms and also to curation criteria. Finally, we found that answer pooling was much faster and allowed us to identify more conflicting genes than interannotator analysis.