Large-scale directional relationship extraction and resolution.

Large-scale directional relationship extraction and resolution.
复制标题

DOI:
10.1186/1471-2105-9-s9-s11
复制
发表时间:
2008-08-12
期刊:
影响因子:
3
通讯作者:
Wren JD
Wren JD
中科院分区:
生物学4区
文献类型:
--
作者:
Giles CB;Wren JD

文献摘要

被引文献

相似文献

MEDLINE中的基因、化学物质、代谢产物、表型和疾病等实体之间的关系往往是方向性的。也就是说,一方可能以积极或消极的方式影响另一方。检测因果关系和方向是拼接路径和检验实验结果的可能含义的关键。由于生物医学文献的规模和增长,能够尽可能地自动化这一过程变得越来越重要。本文提出了一种基于依赖图分析和支持向量机分类的关系抽取方法。我们首先在Genia的黄金标准语料库上测试了支持向量机分类器,在这些标准化测试集上达到了82%的准确率和94.8%的召回率(F-MEASURE为87.9)。然后,我们将整个系统应用于所有可用的MEDLINE摘要,用于两个具有已知效果的目标交互。我们发现,虽然一些方向关系是以低歧义提取的,但其他方向关系显然是矛盾的,至少在孤立的背景下是这样的。当仔细研究时,很明显,其中一些取决于周围的环境(例如,这种关系是指短期还是长期影响,或者焦点是细胞外还是细胞内)。基于叙词表的方向关系抽取可以达到较高的准确率,但在较大的语料库上,由于名词修饰语的影响,容易出现误报。此外,解决或消除关系上下文和偶发事件的方法对于大规模语料库来说是重要的。
Relationships between entities such as genes, chemicals, metabolites, phenotypes and diseases in MEDLINE are often directional. That is, one may affect the other in a positive or negative manner. Detection of causality and direction is key in piecing pathways together and in examining possible implications of experimental results. Because of the size and growth of biomedical literature, it is increasingly important to be able to automate this process as much as possible. Here we present a method of relation extraction using dependency graph parsing with SVM classification. We tested the SVM classifier first on gold standard corpora from GENIA and find it achieved 82% precision and 94.8% recall (F-measure of 87.9) on these standardized test sets. We then applied the entire system to all available MEDLINE abstracts for two target interactions with known effects. We find that while some directional relations are extracted with low ambiguity, others are apparently contradictory, at least when considered in an isolated context. When examined, it is apparent some are dependent upon the surrounding context (e.g. whether the relationship referred to short-term or long-term effects, or whether the focus was extracellular versus intracellular). Thesaurus-based directional relation extraction can be done with reasonable accuracy, but is prone to false-positives on larger corpora due to noun modifiers. Furthermore, methods of resolving or disambiguating relationship context and contingencies are important for large-scale corpora.