Algorithms and semantic infrastructure for mutation impact extraction and grounding.

Algorithms and semantic infrastructure for mutation impact extraction and grounding.
复制标题

DOI:
10.1186/1471-2164-11-s4-s24
复制
发表时间:
2010-12-02
期刊:
影响因子:
4.4
通讯作者:
Baker CJ
Baker CJ
中科院分区:
生物学2区
文献类型:
--
作者:
Laurila JB;Naderi N;Witte R;Riazanov A;Kouznetsov A;Baker CJ

文献摘要

被引文献

相似文献

突变影响提取是目前最先进的突变提取系统中尚未完成的任务。蛋白质突变及其对蛋白质特性的影响隐藏在科学文献中,这使得蛋白质工程师很难获得它们,而表型预测系统目前依赖于人工管理的基因组变异数据库。我们提出了第一个基于规则的方法来提取突变对蛋白质特性的影响,将它们的方向性分为积极的、消极的或中性的。此外,提到的蛋白质和突变是基于它们各自的UniProtKB id和选定的蛋白质特性,即蛋白质功能在基因本体中发现的概念。提取的实体被填充到OWL-DL突变影响本体中,方便使用SPARQL对突变影响进行复杂查询。我们举例说明检索蛋白质和突变序列对特定蛋白质特性的特定方向的影响。此外,我们通过使用SADI(语义自动发现和集成)框架的语义web服务提供对数据的程序化访问。我们通过创建新的突变影响提取方法来解决以非结构化形式访问遗留突变数据的问题,这些方法在由领域专家标记的关于卤代烷脱卤酶的全文文章语料库上进行评估。我们的方法显示了突变基础的精度和召回率的最新水平,并且在突变影响关系提取的任务中具有可观的精度水平,但召回率较低。该系统使用文本挖掘和语义web技术进行部署,目标是向广泛的消费者发布。
Mutation impact extraction is a hitherto unaccomplished task in state of the art mutation extraction systems. Protein mutations and their impacts on protein properties are hidden in scientific literature, making them poorly accessible for protein engineers and inaccessible for phenotype-prediction systems that currently depend on manually curated genomic variation databases. We present the first rule-based approach for the extraction of mutation impacts on protein properties, categorizing their directionality as positive, negative or neutral. Furthermore protein and mutation mentions are grounded to their respective UniProtKB IDs and selected protein properties, namely protein functions to concepts found in the Gene Ontology. The extracted entities are populated to an OWL-DL Mutation Impact ontology facilitating complex querying for mutation impacts using SPARQL. We illustrate retrieval of proteins and mutant sequences for a given direction of impact on specific protein properties. Moreover we provide programmatic access to the data through semantic web services using the SADI (Semantic Automated Discovery and Integration) framework. We address the problem of access to legacy mutation data in unstructured form through the creation of novel mutation impact extraction methods which are evaluated on a corpus of full-text articles on haloalkane dehalogenases, tagged by domain experts. Our approaches show state of the art levels of precision and recall for Mutation Grounding and respectable level of precision but lower recall for the task of Mutant-Impact relation extraction. The system is deployed using text mining and semantic web technologies with the goal of publishing to a broad spectrum of consumers.