Large scale application of neural network based semantic role labeling for automated relation extraction from biomedical texts.

Large scale application of neural network based semantic role labeling for automated relation extraction from biomedical texts.
复制标题

DOI:
10.1371/journal.pone.0006393
复制
发表时间:
2009-07-28
期刊:
影响因子:
3.7
通讯作者:
Stümpflen V
Stümpflen V
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Barnickel T;Weston J;Collobert R;Mewes HW;Stümpflen V

文献摘要

被引文献

相似文献

为了减少生命科学中日益增长的文献搜索所花费的时间,已经开发了几种自动提取知识的方法。基于共现的方法可以在可接受的时间内处理像MEDLINE这样的大型文本语料库,但不能提取任何特定类型的语义关系。另一方面,基于句法树的语义关系提取方法计算量大,生成的语义关系树难以解释。生物医学领域的几种自然语言处理(NLP)方法特别侧重于检测一组有限的关系类型。对于系统生物学来说,需要用于检测多种关系类型的通用方法,这些方法还能够处理大型文本语料库,但满足这两个要求的系统数量非常有限。本文介绍了一种快速、准确的基于神经网络的语义角色标注(SRL)程序--SENA(“语义抽取使用神经网络结构”),用于大规模提取生物医学文献中的语义关系。将SENA与生物医学领域中使用的其他SRL系统或句法分析器的处理时间进行比较,发现SENA是目前可用的符合命题银行(PropBank)的最快的SRL程序。在三天内用番泻叶在百个节点集群上标注了百万条生物医学语句。在两个标注句子的测试集上测试了该关系提取方法的准确率,准确率/召回率为0.71/0.43。实验结果表明,本文提出的语义关系抽取方法的准确率和处理速度足以满足大规模应用于生物医学文本的要求。该方法对于所支持的关系类型具有高度的通用性,特别适合于通用、大规模的文本挖掘系统。该方法弥补了缺乏语义关系的基于共现的快速方法与高度专门化和计算量大的自然语言处理方法之间的差距。
To reduce the increasing amount of time spent on literature search in the life sciences, several methods for automated knowledge extraction have been developed. Co-occurrence based approaches can deal with large text corpora like MEDLINE in an acceptable time but are not able to extract any specific type of semantic relation. Semantic relation extraction methods based on syntax trees, on the other hand, are computationally expensive and the interpretation of the generated trees is difficult. Several natural language processing (NLP) approaches for the biomedical domain exist focusing specifically on the detection of a limited set of relation types. For systems biology, generic approaches for the detection of a multitude of relation types which in addition are able to process large text corpora are needed but the number of systems meeting both requirements is very limited. We introduce the use of SENNA (“Semantic Extraction using a Neural Network Architecture”), a fast and accurate neural network based Semantic Role Labeling (SRL) program, for the large scale extraction of semantic relations from the biomedical literature. A comparison of processing times of SENNA and other SRL systems or syntactical parsers used in the biomedical domain revealed that SENNA is the fastest Proposition Bank (PropBank) conforming SRL program currently available. 89 million biomedical sentences were tagged with SENNA on a 100 node cluster within three days. The accuracy of the presented relation extraction approach was evaluated on two test sets of annotated sentences resulting in precision/recall values of 0.71/0.43. We show that the accuracy as well as processing speed of the proposed semantic relation extraction approach is sufficient for its large scale application on biomedical text. The proposed approach is highly generalizable regarding the supported relation types and appears to be especially suited for general-purpose, broad-scale text mining systems. The presented approach bridges the gap between fast, cooccurrence-based approaches lacking semantic relations and highly specialized and computationally demanding NLP approaches.