Mining literature for protein-protein interactions

Mining literature for protein-protein interactions
复制标题

DOI:
10.1093/bioinformatics/17.4.359
复制
发表时间:
2001-04-01
期刊:
影响因子:
5.8
通讯作者:
Eisenberg, D
Eisenberg, D
中科院分区:
生物学3区
文献类型:
--
作者:
Marcotte, EM;Xenarios, I;Eisenberg, D

文献摘要

被引文献

相似文献

动机:生物信息学的一个核心问题是如何以适合计算机分析的形式从当前大量的科学文献中获取信息。我们解决了蛋白质-蛋白质相互作用信息的特殊情况,并表明Medline摘要中单词的频率可以用来确定给定论文是否讨论蛋白质-蛋白质相互作用。对于那些决定讨论这个主题的论文,相关信息可以被捕获到相互作用蛋白质数据库。此外,还可以捕获合适的基因注释。结果:我们的贝叶斯方法根据摘要中发现的判别词的频率,对Medline摘要讨论感兴趣主题的概率进行评分。从260篇Medline摘要的训练集中确定了80多个判别词(例如,complex, interaction, two-hybrid),这些词对应于先前在相互作用蛋白数据库中验证的条目,使用这些词和对数似然评分函数,类似于2000篇Medline摘要被识别为描述酵母蛋白之间的相互作用。这种方法现在形成了相互作用蛋白质数据库快速扩展的基础。
Motivation: A central problem in bioinformatics is how to capture information from the vast current scientific literature in a form suitable for analysis by computer. We address the special case of information on protein-protein interactions, and show that the frequencies of words in Medline abstracts can be used to determine whether or not a given paper discusses protein-protein interactions. For those papers determined to discuss this topic, the relevant information can be captured for the Database of Interacting Proteins. Furthermore, suitable gene annotations can also be captured.Results: Our Bayesian approach scores Medline abstracts for probability of discussing the topic of interest according to the frequencies of discriminating words found in the abstract. More than 80 discriminating words (e.g, complex, interaction, two-hybrid) were determined from a training set of 260 Medline abstracts corresponding to previously validated entries in the Database of Interacting Proteins, Using these words and a log likelihood scoring function, similar to 2000 Medline abstracts were identified as describing interactions between yeast proteins. This approach now forms the basis for the rapid expansion of the Database of Interacting Proteins.