Automatic extraction of protein point mutations using a graph bigram association.

Automatic extraction of protein point mutations using a graph bigram association.
复制标题

DOI:
10.1371/journal.pcbi.0030016
复制
发表时间:
2007-02-02
影响因子:
4.3
通讯作者:
Cohen FE
Cohen FE
中科院分区:
生物学2区
文献类型:
--
作者:
Lee LC;Horn F;Cohen FE

文献摘要

参考文献

被引文献

相似文献

蛋白质点突变是蛋白质结构和功能的进化和实验分析的重要组成部分。虽然许多手动管理的数据库试图索引点突变,但大多数实验产生的点突变和变化的生物学影响在同行评审的已发表文献中进行了描述。我们描述了一个应用程序,突变GraB(图Bigram),识别,提取和验证点突变的生物医学文献。点突变提取的主要问题是将点突变与其相关蛋白质和起源生物联系起来。我们的算法使用基于图的二元遍历来识别这些相关的关联,并利用Swiss-Prot蛋白质数据库来验证这些信息。图二元组方法与其他点突变提取模型的不同之处在于,它结合了文章中所有术语的频率和位置数据来驱动点突变-蛋白质关联。我们的方法在589篇描述G蛋白偶联受体(GPCR)、酪氨酸激酶和离子通道蛋白家族点突变的文章中进行了测试。我们在这三个不同蛋白质家族的全文文献数据集上评估了我们的图二元组度量与词邻近度度量的术语关联。我们的测试表明,图二元度量实现了更高的F-测量GPCR(0.79对0.76),蛋白酪氨酸激酶(0.72对0.69),和离子通道转运蛋白(0.76对0.74)。重要的是,在可以将多于一个蛋白质分配给点突变并且需要消歧的情况下,与词距离度量精度0.73相比,图二元组度量精度达到0.84。我们认为,图二元搜索度量是一个显着的改进,比以前的搜索度量点突变提取,并适用于文本挖掘应用程序需要的关联词。在生物学研究中,新信息通常以同行评审的期刊文章的形式呈现。尽管电子数据库管理员尽了最大的努力,但大多数信息仍然只能以文本形式找到,因此无法直接进行计算分析。科学文献中丰富的一种此类信息是蛋白质点突变。我们试图从文献中提取蛋白质点突变的例子,并将它们与标准化蛋白质数据库中的唯一蛋白质名称和物种起源相关联。为此,我们创建了一个应用程序,用于搜索和检索来自出版商的全文文章,识别文章中的点突变术语、蛋白质名称术语和生物体名称术语。我们描述突变GraB,一个应用程序,利用图最短距离搜索与词二元分析,用于找到这些条款之间的显着关联的文本。该图二元搜索度量被发现在识别正确的蛋白质点突变对方面相当有效,并且代表了准确性和广泛适用性之间的良好折衷。该应用程序可以应用于来自蛋白质家族的大量期刊文献,以生成点突变数据库。
Protein point mutations are an essential component of the evolutionary and experimental analysis of protein structure and function. While many manually curated databases attempt to index point mutations, most experimentally generated point mutations and the biological impacts of the changes are described in the peer-reviewed published literature. We describe an application, Mutation GraB (Graph Bigram), that identifies, extracts, and verifies point mutations from biomedical literature. The principal problem of point mutation extraction is to link the point mutation with its associated protein and organism of origin. Our algorithm uses a graph-based bigram traversal to identify these relevant associations and exploits the Swiss-Prot protein database to verify this information. The graph bigram method is different from other models for point mutation extraction in that it incorporates frequency and positional data of all terms in an article to drive the point mutation–protein association. Our method was tested on 589 articles describing point mutations from the G protein–coupled receptor (GPCR), tyrosine kinase, and ion channel protein families. We evaluated our graph bigram metric against a word-proximity metric for term association on datasets of full-text literature in these three different protein families. Our testing shows that the graph bigram metric achieves a higher F-measure for the GPCRs (0.79 versus 0.76), protein tyrosine kinases (0.72 versus 0.69), and ion channel transporters (0.76 versus 0.74). Importantly, in situations where more than one protein can be assigned to a point mutation and disambiguation is required, the graph bigram metric achieves a precision of 0.84 compared with the word distance metric precision of 0.73. We believe the graph bigram search metric to be a significant improvement over previous search metrics for point mutation extraction and to be applicable to text-mining application requiring the association of words. In biological research, new information is often presented in the form of peer-reviewed published journal articles. Despite the best efforts of electronic database curators, a majority of this information is still found only in textual form, and thus excluded from direct computational analysis. One such type of information that is abundant in scientific literature is protein point mutations. We seek to extract protein point mutation examples from the literature and to associate them with a unique protein name and species of origin in a standardized protein database. To do this, we have created an application that searches for and retrieves full-text articles from publishers, identifies point mutation terms, protein name terms, and organism name terms within the articles. We describe Mutation GraB, an application that utilizes a graph shortest-distance search in concert with word bigram analysis that is used to find significant associations between these terms in the text. This graph bigram search metric was found to be reasonably effective at identifying correct protein point mutation pairs and represents a good compromise between accuracy and broad applicability. The application can be applied to a large set of journal literature from a protein family to generate a database of point mutations.
DOI: 10.1093/bioinformatics/btg393
发表时间: 2004-01-22
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Chang, JT;Schütze, H;Altman, RB
通讯作者: Altman, RB
DOI: 10.1038/sj.tpj.6500070
发表时间: 2002-01-01
影响因子: 2.8
作者:
Cotton, R. G. H.;Horaitis, O.
通讯作者: Horaitis, O.
DOI: 10.1093/nar/gkh162
发表时间: 2004-01-01
影响因子: 14.9
作者:
Rebholz-Schuhmann, D;Marcel, S;Kirsch, H
通讯作者: Kirsch, H
DOI: 10.1186/1471-2105-6-s1-s15
发表时间: 2005
期刊: BMC bioinformatics
影响因子: 3
作者:
Fundel K;Güttler D;Zimmer R;Apostolakis J
通讯作者: Apostolakis J
DOI: 10.1186/1471-2105-6-s1-s13
发表时间: 2005
期刊: BMC bioinformatics
影响因子: 3
作者:
Crim J;McDonald R;Pereira F
通讯作者: Pereira F