New directions in biomedical text annotation: definitions, guidelines and corpus construction.

New directions in biomedical text annotation: definitions, guidelines and corpus construction.
复制标题

DOI:
10.1186/1471-2105-7-356
复制
发表时间:
2006-07-25
期刊:
影响因子:
3
通讯作者:
--
中科院分区:
生物学4区
文献类型:
--
作者:

文献摘要

被引文献

相似文献

虽然生物医学文本挖掘正在成为一个重要的研究领域,但实际结果已经证明很难实现。我们认为,更准确的文本挖掘的重要的第一步在于能够识别和描述满足各种类型的信息需求的文本。我们在这里报告的结果,我们的调查性质的科学文本,有足够的一般性,以超越一个狭窄的学科领域的限制,同时支持实际挖掘的文本的事实信息。我们的最终目标是注释一个重要的生物医学文本语料库,并训练机器学习方法来自动将这些文本沿着我们定义的沿着某些维度进行分类。我们已经确定了五个定性维度,我们认为表征了广泛的科学句子,因此有助于支持文本挖掘的一般方法:焦点,极性,确定性,证据和方向性。我们定义了这些维度,并描述了我们为注释文本而制定的准则。为了检验指南的有效性,12名注释者独立地注释了从当前生物医学期刊中随机选择的101个句子。对这些注释的分析表明,注释者之间的一致性为70-80%,这表明我们的指南确实提供了一个定义良好、可执行和可复制的任务。我们提出了我们的指导方针,定义一个文本注释任务,沿着注释结果从多个独立产生的注释,证明了任务的可行性。目前正在根据这些准则对大量文件进行沿着说明。这些注释形成了文本沿着多个维度分类的基础,以支持对实验结果、方法论声明和其他形式的信息的可行文本挖掘。我们目前正在开发机器学习方法,在注释语料库上进行训练和测试,这将允许生物医学文本的自动分类沿着我们已经提出的一般维度。这些准则的全部细节,连同附有注释的例子,沿着可供公众查阅。
While biomedical text mining is emerging as an important research area, practical results have proven difficult to achieve. We believe that an important first step towards more accurate text-mining lies in the ability to identify and characterize text that satisfies various types of information needs. We report here the results of our inquiry into properties of scientific text that have sufficient generality to transcend the confines of a narrow subject area, while supporting practical mining of text for factual information. Our ultimate goal is to annotate a significant corpus of biomedical text and train machine learning methods to automatically categorize such text along certain dimensions that we have defined. We have identified five qualitative dimensions that we believe characterize a broad range of scientific sentences, and are therefore useful for supporting a general approach to text-mining: focus, polarity, certainty, evidence, and directionality. We define these dimensions and describe the guidelines we have developed for annotating text with regard to them. To examine the effectiveness of the guidelines, twelve annotators independently annotated the same set of 101 sentences that were randomly selected from current biomedical periodicals. Analysis of these annotations shows 70–80% inter-annotator agreement, suggesting that our guidelines indeed present a well-defined, executable and reproducible task. We present our guidelines defining a text annotation task, along with annotation results from multiple independently produced annotations, demonstrating the feasibility of the task. The annotation of a very large corpus of documents along these guidelines is currently ongoing. These annotations form the basis for the categorization of text along multiple dimensions, to support viable text mining for experimental results, methodology statements, and other forms of information. We are currently developing machine learning methods, to be trained and tested on the annotated corpus, that would allow for the automatic categorization of biomedical text along the general dimensions that we have presented. The guidelines in full detail, along with annotated examples, are publicly available.