Can bibliographic pointers for known biological data be found automatically? Protein interactions as a case study.

Can bibliographic pointers for known biological data be found automatically? Protein interactions as a case study.
复制标题

DOI:
10.1002/cfg.91
复制
发表时间:
2001
影响因子:
--
通讯作者:
Valencia, A
Valencia, A
中科院分区:
其他
文献类型:
--
作者:
Blaschke, C;Valencia, A

文献摘要

被引文献

相似文献

相互作用蛋白质词典(DIP) (Xenarios et al., 2000)是一个蛋白质相互作用的大型资料库:其2000年3月发布的版本包括2379对蛋白质,它们的相互作用已通过实验方法检测到。即使其中许多与特征不明确的蛋白质相对应,大量酵母双杂交筛选的结果是,多达851个与使用直接生化方法检测到的相互作用相对应。我们使用信息检索技术在Medline摘要中自动搜索支持这些851 DIP交互的句子。令人惊讶的是,我们发现DIP蛋白对和Medline句子之间的对应关系只在30%的情况下描述了它们的相互作用。这种低覆盖率对数据库中引入的注释(参考文献)的质量和信息提取(IE)技术在分子生物学中的应用的局限性产生了有趣的影响。很明显,分析摘要而不是全文的局限性和缺乏标准的蛋白质名称是比IE方法的局限性更重要的困难。一个积极的发现是IE系统能够识别蛋白质之间的新关系,即使是在一组先前由人类专家鉴定的蛋白质中。这些鉴定是相当精确的。据我们所知,这是第一次对IE能力进行大规模评估,以检测先前已知的相互作用:因此,我们建议使用DIP数据集作为基准IE系统的生物学参考。
The Dictionary of Interacting Proteins (DIP) (Xenarios et al., 2000) is a large repository of protein interactions: its March 2000 release included 2379 protein pairs whose interactions have been detected by experimental methods. Even if many of these correspond to poorly characterized proteins, the result of massive yeast two-hybrid screenings, as many as 851 correspond to interactions detected using direct biochemical methods. We used information retrieval technology to search automatically for sentences in Medline abstracts that support these 851 DIP interactions. Surprisingly, we found correspondence between DIP protein pairs and Medline sentences describing their interactions in only 30% of the cases. This low coverage has interesting consequences regarding the quality of annotations (references) introduced in the database and the limitations of the application of information extraction (IE) technology to Molecular Biology. It is clear that the limitation of analyzing abstracts rather than full papers and the lack of standard protein names are difficulties of considerably more importance than the limitations of the IE methodology employed. A positive finding is the capacity of the IE system to identify new relations between proteins, even in a set of proteins previously characterized by human experts. These identifications are made with a considerable degree of precision. This is, to our knowledge, the first large scale assessment of IE capacity to detect previously known interactions: we thus propose the use of the DIP data set as a biological reference to benchmark IE systems.