Gene mention normalization and interaction extraction with context models and sentence motifs.

Gene mention normalization and interaction extraction with context models and sentence motifs.
复制标题

基因提到了与上下文模型和句子图案的归一化和互动提取。

DOI:
10.1186/gb-2008-9-s2-s14
复制
发表时间:
2008
期刊:
影响因子:
12.3
通讯作者:
Schroeder M
Schroeder M
中科院分区:
生物学1区
文献类型:
--
作者:
Hakenberg J;Plake C;Royer L;Strobelt H;Leser U;Schroeder M

文献摘要

参考文献

被引文献

相似文献

文本挖掘的目标是使科学出版物中传达的信息可用于结构化搜索和自动分析。文本挖掘的两个重要子任务是实体提及规范化-识别文本中的生物医学对象-以及提取这些对象之间的合格关系。我们描述了一种用于识别基因和蛋白质之间的关系的方法。我们提出了解决方案,基因提及规范化和提取蛋白质-蛋白质相互作用。对于第一个任务,我们通过使用每个基因的背景知识来识别基因,即与功能,位置,疾病等相关的注释。我们的方法目前在BioCreative II基因标准化数据上实现了86.4%的f-测量。对于蛋白质-蛋白质相互作用的提取,我们追求一种建立在经典序列分析基础上的方法:来自多序列比对的基序。该方法在BioCreative II交互对子任务中实现了24.4%(微平均值)的f测量。对于基因提及规范化,我们的方法优于仅利用基因名称与字典匹配的策略,而无需调用每个基因的进一步知识。来自句子对齐的基序在识别文本中的蛋白质相互作用方面是成功的;我们在本报告中提出的方法是完全自动化的,并且与在一个或多个阶段需要人工干预的系统类似。我们的基因、蛋白质和物种鉴定以及蛋白质-蛋白质提取方法可作为BioCreative Meta Services(BCMS)的一部分,参见。
The goal of text mining is to make the information conveyed in scientific publications accessible to structured search and automatic analysis. Two important subtasks of text mining are entity mention normalization - to identify biomedical objects in text - and extraction of qualified relationships between those objects. We describe a method for identifying genes and relationships between proteins. We present solutions to gene mention normalization and extraction of protein-protein interactions. For the first task, we identify genes by using background knowledge on each gene, namely annotations related to function, location, disease, and so on. Our approach currently achieves an f-measure of 86.4% on the BioCreative II gene normalization data. For the extraction of protein-protein interactions, we pursue an approach that builds on classical sequence analysis: motifs derived from multiple sequence alignments. The method achieves an f-measure of 24.4% (micro-average) in the BioCreative II interaction pair subtask. For gene mention normalization, our approach outperforms strategies that utilize only the matching of genes names against dictionaries, without invoking further knowledge on each gene. Motifs derived from alignments of sentences are successful at identifying protein interactions in text; the approach we present in this report is fully automated and performs similarly to systems that require human intervention at one or more stages. Our methods for gene, protein, and species identification, and extraction of protein-protein are available as part of the BioCreative Meta Services (BCMS), see .
DOI: 10.1109/mis.2002.999215
发表时间: 2002-03-01
影响因子: 6.4
作者:
Blaschke, C;Valencia, A
通讯作者: Valencia, A
DOI: 10.1093/bioinformatics/bti597
发表时间: 2006-03-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Saric, J;Jensen, LJ;Bork, P
通讯作者: Bork, P
DOI: 10.1093/bioinformatics/btl408
发表时间: 2006-10-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Plake, Conrad;Schiemann, Torsten;Leser, Ulf
通讯作者: Leser, Ulf
DOI: 10.1038/nature04532
发表时间: 2006-03-30
期刊: NATURE
影响因子: 64.8
作者:
Gavin, AC;Aloy, P;Superti-Furga, G
通讯作者: Superti-Furga, G
评估生物学的文本挖掘系统:第二次生物综合社区挑战的概述。
DOI: 10.1186/gb-2008-9-s2-s1
发表时间: 2008
期刊: Genome biology
影响因子: 12.3
作者:
Krallinger M;Morgan A;Smith L;Leitner F;Tanabe L;Wilbur J;Hirschman L;Valencia A
通讯作者: Valencia A