Structured learning for spatial information extraction from biomedical text: bacteria biotopes.

Structured learning for spatial information extraction from biomedical text: bacteria biotopes.
复制标题

DOI:
10.1186/s12859-015-0542-z
复制
发表时间:
2015-04-25
期刊:
影响因子:
3
通讯作者:
Moens MF
Moens MF
中科院分区:
生物学4区
文献类型:
--
作者:
Kordjamshidi P;Roth D;Moens MF

文献摘要

参考文献

被引文献

相似文献

我们的目标是从网页中自动提取细菌的物种名称和它们的位置。这项任务对于利用以不同自然语言文本表达的大量生物学知识并将这些知识放入数据库以便于生物学家访问非常重要。这项任务是具有挑战性的,以前的结果是远远低于可接受的性能水平,特别是对于本地化关系的提取。因此,我们的目标是设计一个新的系统,这样的提取,使用结构化机器学习技术的框架。设计了一种新的生物医学实体联合提取模型和定位关系。我们的模型是基于一个空间的角色标记(SpRL)模型设计的空间理解不受限制的文本。我们扩展了SpRL,以提取生物医学领域的话语级空间关系,并将其应用于BioNLP-ST 2013,BB共享任务。我们强调一般的空间语言理解和空间信息提取的科学文本,这是这项工作的重点之间的主要区别。我们利用文本的结构和语篇层次的全球功能。我们的模型和设计的功能大大改善了以前的系统,实现了约57%的绝对改善超过F1的最佳以前的系统为这项任务的措施。我们的实验结果表明,在一个文档中的所有实体和关系的联合学习模型优于一个模型,独立提取实体和关系。我们的全局学习模型显着提高了这项任务的最先进的结果,并有很大的潜力被采用在其他自然语言处理(NLP)任务在生物医学领域。
We aim to automatically extract species names of bacteria and their locations from webpages. This task is important for exploiting the vast amount of biological knowledge which is expressed in diverse natural language texts and putting this knowledge in databases for easy access by biologists. The task is challenging and the previous results are far below an acceptable level of performance, particularly for extraction of localization relationships. Therefore, we aim to design a new system for such extractions, using the framework of structured machine learning techniques. We design a new model for joint extraction of biomedical entities and the localization relationship. Our model is based on a spatial role labeling (SpRL) model designed for spatial understanding of unrestricted text. We extend SpRL to extract discourse level spatial relations in the biomedical domain and apply it on the BioNLP-ST 2013, BB-shared task. We highlight the main differences between general spatial language understanding and spatial information extraction from the scientific text which is the focus of this work. We exploit the text’s structure and discourse level global features. Our model and the designed features substantially improve on the previous systems, achieving an absolute improvement of approximately 57 percent over F1 measure of the best previous system for this task. Our experimental results indicate that a joint learning model over all entities and relationships in a document outperforms a model which extracts entities and relationships independently. Our global learning model significantly improves the state-of-the-art results on this task and has a high potential to be adopted in other natural language processing (NLP) tasks in the biomedical domain.
DOI: 10.1093/nar/gki051
发表时间: 2005-01-01
影响因子: 14.9
作者:
Alfarano C;Andrade CE;Anthony K;Bahroos N;Bajec M;Bantoft K;Betel D;Bobechko B;Boutilier K;Burgess E;Buzadzija K;Cavero R;D'Abreo C;Donaldson I;Dorairajoo D;Dumontier MJ;Dumontier MR;Earles V;Farrall R;Feldman H;Garderman E;Gong Y;Gonzaga R;Grytsan V;Gryz E;Gu V;Haldorsen E;Halupa A;Haw R;Hrvojic A;Hurrell L;Isserlin R;Jack F;Juma F;Khan A;Kon T;Konopinsky S;Le V;Lee E;Ling S;Magidin M;Moniakis J;Montojo J;Moore S;Muskat B;Ng I;Paraiso JP;Parker B;Pintilie G;Pirone R;Salama JJ;Sgro S;Shan T;Shu Y;Siew J;Skinner D;Snyder K;Stasiuk R;Strumpf D;Tuekam B;Tao S;Wang Z;White M;Willis R;Wolting C;Wong S;Wrong A;Xin C;Yao R;Yates B;Zhang S;Zheng K;Pawson T;Ouellette BF;Hogue CW
通讯作者: Hogue CW
DOI: 10.1186/2041-1480-3-3
发表时间: 2012-04-01
影响因子: 1.9
作者:
Liu H;Christiansen T;Baumgartner WA Jr;Verspoor K
通讯作者: Verspoor K
DOI: 10.1016/j.websem.2014.06.001
发表时间: 2015-01-01
影响因子: 2.5
作者:
Kordjamshidi, Parisa;Moens, Marie-Francine
通讯作者: Moens, Marie-Francine
DOI: 10.1162/jmlr.2003.3.4-5.679
发表时间: 2003-05-15
影响因子: 6
作者:
Getoor, L;Friedman, N;Taskar, B
通讯作者: Taskar, B
DOI: 10.1021/bi00231a020
发表时间: 1991-04-30
期刊: BIOCHEMISTRY
影响因子: 2.9
作者:
MERUTKA, G;SHALONGO, W;STELLWAGEN, E
通讯作者: STELLWAGEN, E