Extraction of human kinase mutations from literature, databases and genotyping studies.

Extraction of human kinase mutations from literature, databases and genotyping studies.
复制标题

DOI:
10.1186/1471-2105-10-s8-s1
复制
发表时间:
2009-08-27
期刊:
影响因子:
3
通讯作者:
Valencia A
Valencia A
中科院分区:
生物学4区
文献类型:
--
作者:
Krallinger M;Izarzugaza JM;Rodriguez-Penagos C;Valencia A

文献摘要

被引文献

相似文献

通过诱变实验来表征特定蛋白质残基取代的生物学作用具有相当大的兴趣。此外,最近与疾病相关的SNP检测相关的努力激发了从文献中手动注释以及自动提取天然存在的序列变异,特别是对于在信号传导过程中起重要作用的蛋白质家族,如激酶。系统整合和比较来自多个来源的激酶突变信息,包括文献、手动注释数据库和大规模实验,可以更全面地了解蛋白质序列变体的功能、结构和疾病相关方面。以前发表的突变提取方法没有充分区分两种根本不同的变异起源类别,即通过体外实验产生的天然发生的和诱导的突变。我们提出了一个文献挖掘管道,用于自动提取和消除摘要和全文文章中提到的单点突变的歧义,然后进行序列验证检查,将突变与其相应的激酶蛋白序列联系起来。每个突变根据其是否对应于诱导突变或天然序列变体来评分。我们能够为相当一部分先前注释的激酶突变提供直接的文献链接,从而能够更有效地解释其生物学特性和实验背景。为了测试所提出的流水线的能力,分析了激酶家族的蛋白激酶结构域中的突变。使用我们的文献提取系统,我们能够从PubMed摘要中恢复总共643个突变-蛋白质关联,并从大量的全文文章中恢复6,970个。当与最先进的注释数据库和高通量基因分型研究相比时,从文献中提取的突变提及与现有知识库在很大程度上重叠,而其余提及表明先前未在数据库中注释的新突变记录。使用建议的残基消歧和分类方法,我们能够区分自然变异和突变类型的突变,准确率为93.88。由此产生的系统是有用的构建黄金标准集的突变从文献中提取的人类专家以最小的手动策展工作,提供直接指针相关的证据句子。我们的系统能够从文献中恢复突变,这些突变在最先进的数据库中不存在。对PubMed摘要中100个突变进行的文献提取突变子集的人类专家手动验证强调,几乎四分之三(72%)的提取突变被证明是正确的,其中一半以上以前没有在数据库中注释。
There is a considerable interest in characterizing the biological role of specific protein residue substitutions through mutagenesis experiments. Additionally, recent efforts related to the detection of disease-associated SNPs motivated both the manual annotation, as well as the automatic extraction, of naturally occurring sequence variations from the literature, especially for protein families that play a significant role in signaling processes such as kinases. Systematic integration and comparison of kinase mutation information from multiple sources, covering literature, manual annotation databases and large-scale experiments can result in a more comprehensive view of functional, structural and disease associated aspects of protein sequence variants. Previously published mutation extraction approaches did not sufficiently distinguish between two fundamentally different variation origin categories, namely natural occurring and induced mutations generated through in vitro experiments. We present a literature mining pipeline for the automatic extraction and disambiguation of single-point mutation mentions from both abstracts as well as full text articles, followed by a sequence validation check to link mutations to their corresponding kinase protein sequences. Each mutation is scored according to whether it corresponds to an induced mutation or a natural sequence variant. We were able to provide direct literature links for a considerable fraction of previously annotated kinase mutations, enabling thus more efficient interpretation of their biological characterization and experimental context. In order to test the capabilities of the presented pipeline, the mutations in the protein kinase domain of the kinase family were analyzed. Using our literature extraction system, we were able to recover a total of 643 mutations-protein associations from PubMed abstracts and 6,970 from a large collection of full text articles. When compared to state-of-the-art annotation databases and high throughput genotyping studies, the mutation mentions extracted from the literature overlap to a good extent with the existing knowledgebases, whereas the remaining mentions suggest new mutation records that were not previously annotated in the databases. Using the proposed residue disambiguation and classification approach, we were able to differentiate between natural variant and mutagenesis types of mutations with an accuracy of 93.88. The resulting system is useful for constructing a Gold Standard set of mutations extracted from the literature by human experts with minimal manual curation effort, providing direct pointers to relevant evidence sentences. Our system is able to recover mutations from the literature that are not present in state-of-the-art databases. Human expert manual validation of a subset of the literature extracted mutations conducted on 100 mutations from PubMed abstracts highlights that almost three quarters (72%) of the extracted mutations turned out to be correct, and more than half of these had not been previously annotated in databases.