Extracting human protein interactions from MEDLINE using a full-sentence parser

Extracting human protein interactions from MEDLINE using a full-sentence parser
复制标题

DOI:
10.1093/bioinformatics/btg452
复制
发表时间:
2004-03-22
期刊:
影响因子:
5.8
通讯作者:
Mazo, I
Mazo, I
中科院分区:
生物学3区
文献类型:
--
作者:
Daraselia, N;Yuryev, A;Mazo, I

文献摘要

被引文献

相似文献

动机:活细胞是一个复杂的机器,它依赖于它的许多部分的正常运作,包括蛋白质。了解蛋白质功能以及它们如何相互修饰和调节是生命科学研究人员面临的下一个重大挑战。有关蛋白质功能和途径的集体知识分散在科学期刊的众多出版物中。将相关信息汇集在一起成为研究和发现过程中的瓶颈。这些信息的数量呈指数级增长,这使得手动管理变得不切实际。作为一个可行的替代方案,自动化的文献处理工具可以用来提取和组织生物数据到一个知识库,使其适合计算分析和数据mining.Results:我们提出MedScan,一个完全自动化的基于自然语言处理的信息提取系统。我们使用MedScan从1988年以后的MEDLINE摘要中提取了2976种人类蛋白质之间的相互作用。提取的信息的精度被认为是91%。与现有的蛋白质相互作用数据库BIND和DIP的比较表明,96%的提取信息是新的。MedScan的召回率为21%。MedScan的其他实验表明,MEDLINE是不同蛋白质功能信息的独特来源,可以以相当高的精度以完全自动化的方式提取。进一步的MedScan技术改进的方向进行了讨论。
Motivation: The living cell is a complex machine that depends on the proper functioning of its numerous parts, including proteins. Understanding protein functions and how they modify and regulate each other is the next great challenge for life-sciences researchers. The collective knowledge about protein functions and pathways is scattered throughout numerous publications in scientific journals. Bringing the relevant information together becomes a bottleneck in a research and discovery process. The volume of such information grows exponentially, which renders manual curation impractical. As a viable alternative, automated literature processing tools could be employed to extract and organize biological data into a knowledge base, making it amenable to computational analysis and data mining.Results: We present MedScan, a completely automated natural language processing-based information extraction system. We have used MedScan to extract 2976 interactions between human proteins from MEDLINE abstracts dated after 1988. The precision of the extracted information was found to be 91%. Comparison with the existing protein interaction databases BIND and DIP revealed that 96% of extracted information is novel. The recall rate of MedScan was found to be 21%. Additional experiments with MedScan suggest that MEDLINE is a unique source of diverse protein function information, which can be extracted in a completely automated way with a reasonably high precision. Further directions of the MedScan technology improvement are discussed.