A survey on annotation tools for the biomedical literature

A survey on annotation tools for the biomedical literature
复制标题

DOI:
10.1093/bib/bbs084
复制
发表时间:
2014-03-01
影响因子:
9.5
通讯作者:
Leser, Ulf
Leser, Ulf
中科院分区:
生物学2区
文献类型:
--
作者:
Neves, Mariana;Leser, Ulf

文献摘要

被引文献

相似文献

生物医学文本挖掘的新方法关键依赖于全面的标注语料库的存在。这种语料库通常被称为黄金标准,对于培训阶段的学习模式或模型、评估和比较算法的性能以及更好地理解通过例子寻找的信息都很重要。黄金标准依赖于人类对自然语言文本的理解和手动注释。这个过程非常耗时和昂贵,因为它需要领域专家的高度智力努力。因此,缺乏黄金标准被认为是开发新的文本挖掘方法的主要瓶颈之一。这种情况导致了支持人类注释文本的工具的开发。这样的工具应该直观易用,应该支持一系列不同的输入格式,应该包括注释文本的可视化,并且应该生成易于解析的输出格式。今天,实现其中一些功能的一系列工具都可用。在这次调查中,我们提出了一个全面的调查工具,以支持生物医学文本的注释。我们总共考虑了近30种工具,其中13种被选中进行深入比较。比较是使用预定义的标准进行的,并尽可能伴随着实践经验。我们的调查表明,目前的工具能够以令人满意的方式支持生物医学文本标注中的许多任务,但也没有工具可以被认为是真正全面的解决方案。
New approaches to biomedical text mining crucially depend on the existence of comprehensive annotated corpora. Such corpora, commonly called gold standards, are important for learning patterns or models during the training phase, for evaluating and comparing the performance of algorithms and also for better understanding the information sought for by means of examples. Gold standards depend on human understanding and manual annotation of natural language text. This process is very time-consuming and expensive because it requires high intellectual effort from domain experts. Accordingly, the lack of gold standards is considered as one of the main bottlenecks for developing novel text mining methods. This situation led the development of tools that support humans in annotating texts. Such tools should be intuitive to use, should support a range of different input formats, should include visualization of annotated texts and should generate an easy-to-parse output format. Today, a range of tools which implement some of these functionalities are available. In this survey, we present a comprehensive survey of tools for supporting annotation of biomedical texts. Altogether, we considered almost 30 tools, 13 of which were selected for an in-depth comparison. The comparison was performed using predefined criteria and was accompanied by hands-on experiences whenever possible. Our survey shows that current tools can support many of the tasks in biomedical text annotation in a satisfying manner, but also that no tool can be considered as a true comprehensive solution.