Assignment of protein function and discovery of novel nucleolar proteins based on automatic analysis of MEDLINE

Assignment of protein function and discovery of novel nucleolar proteins based on automatic analysis of MEDLINE
复制标题

DOI:
10.1002/pmic.200600693
复制
发表时间:
2007-03-01
期刊:
影响因子:
3.4
通讯作者:
Mons, Barend
Mons, Barend
中科院分区:
生物学3区
文献类型:
--
作者:
Schuemie, Martijn;Chichester, Christine;Mons, Barend

文献摘要

被引文献

相似文献

最可能的功能归因于蛋白质组学确定的蛋白质是一个重大挑战,需要广泛的文献分析。我们已经开发了一个用于自动预测隐式和显式生物学上有意义的功能的系统,以用于核仁的蛋白质组学研究。该方法使用一组词汇术语来映射和集成整个MEDLINE数据库中的信息。基于跨物种序列同源性搜索和相应文献的组合,我们的方法促进了序列数据与描述功能的生物学文本中信息之间的直接关联。将自动化功能分配与手动注释的比较证明了我们的方法非常有效。为了建立灵敏度,我们定义了包含高度保守序列的家庭中的功能微妙。 Dead-box蛋白RNA解旋酶家族的聚类证实,这些蛋白质具有相似的形态,尽管通过我们的方法准确地鉴定了功能亚家族。我们使用多维缩放尺度以蛋白质功能来观察核仁蛋白质组,显示了以前未实现的核仁蛋白之间的功能关联。最后,通过聚集已建立的核仁蛋白的功能特性,我们预测了新型的核仁蛋白。随后,非蛋白质组学研究证实了先前未鉴定的核仁蛋白的预测。
Attribution of the most probable functions to proteins identified by proteomics is a significant challenge that requires extensive literature analysis. We have developed a system for automated prediction of implicit and explicit biologically meaningful functions for a proteomics study of the nucleolus. This approach uses a set of vocabulary terms to map and integrate the information from the entire MEDLINE database. Based on a combination of cross-species sequence homology searches and the corresponding literature, our approach facilitated the direct association between sequence data and information from biological texts describing function. Comparison of our automated functional assignment to manual annotation demonstrated our method to be highly effective. To establish the sensitivity, we defined the functional subtleties within a family containing a highly conserved sequence. Clustering of the DEAD-box protein family of RNA helicases confirmed that these proteins shared similar morphology although functional subfamilies were accurately identified by our approach. We visualized the nucleolar proteome in terms of protein functions using multi-dimensional scaling, showing functional associations between nucleolar proteins that were not previously realized. Finally, by clustering the functional properties of the established nucleolar proteins, we predicted novel nucleolar proteins. Subsequently, nonproteomics studies confirmed the predictions of previously unidentified nucleolar proteins.