An integrative computational model for large-scale identification of metalloproteins in microbial genomes: a focus on iron-sulfur cluster proteins

An integrative computational model for large-scale identification of metalloproteins in microbial genomes: a focus on iron-sulfur cluster proteins
复制标题

DOI:
10.1039/c4mt00156g
复制
发表时间:
2014-01-01
期刊:
影响因子:
3.4
通讯作者:
Vandenbrouck, Yves
Vandenbrouck, Yves
中科院分区:
生物学2区
文献类型:
--
作者:
Estellon, Johan;de Choudens, Sandrine Ollagnier;Vandenbrouck, Yves

文献摘要

被引文献

相似文献

金属蛋白是一组普遍存在的分子,对所有生物体的生存至关重要。虽然已经定义了几种金属结合基序,但仅使用计算方法从主要蛋白质序列中确定金属蛋白仍然具有挑战性。在这里,我们描述了一种基于机器学习方法的综合策略,用于设计和评估惩罚广义线性模型。我们使用这种策略来检测铁硫簇蛋白家族的成员。一种新的描述子,其轮廓是基于轮廓隐马尔可夫模型,编码结构信息与公共描述子组合成一个线性模型。该模型在由充分表征的铁硫蛋白序列组成的不同数据集上进行了训练和测试,与基于基序的方法相比,所得模型提供了更高的灵敏度,同时保持了良好的特异性水平。这个线性模型的分析使我们能够检测和量化每个描述符的贡献,为我们提供了一个更好的理解这个复杂的蛋白质家族沿着有价值的迹象,进一步的实验表征。两个新鉴定的蛋白质YhcC和YdiJ在功能上被验证为真正的铁硫蛋白,证实了预测。然后将计算模型应用于超过550个原核基因组以筛选铁硫蛋白质组;结果可在http://biodev.extra.cea.fr/isph上公开获得。这项研究是一个概念验证的应用惩罚线性模型,以确定金属蛋白超家族的大规模。这里采用的应用程序,筛选铁硫蛋白质组,提供了新的候选人,进一步的生化和结构分析,以及新的资源,广泛探索铁硫蛋白在微生物世界。
Metalloproteins represent a ubiquitous group of molecules which are crucial to the survival of all living organisms. While several metal-binding motifs have been defined, it remains challenging to confidently identify metalloproteins from primary protein sequences using computational approaches alone. Here, we describe a comprehensive strategy based on a machine learning approach to design and assess a penalized generalized linear model. We used this strategy to detect members of the iron-sulfur cluster protein family. A new category of descriptors, whose profile is based on profile hidden Markov models, encoding structural information was combined with public descriptors into a linear model. The model was trained and tested on distinct datasets composed of well-characterized iron-sulfur protein sequences, and the resulting model provided higher sensitivity compared to a motif-based approach, while maintaining a good level of specificity. Analysis of this linear model allows us to detect and quantify the contribution of each descriptor, providing us with a better understanding of this complex protein family along with valuable indications for further experimental characterization. Two newly-identified proteins, YhcC and YdiJ, were functionally validated as genuine iron-sulfur proteins, confirming the prediction. The computational model was then applied to over 550 prokaryotic genomes to screen for iron-sulfur proteomes; the results are publicly available at: http://biodev.extra.cea.fr/isph. This study represents a proof-of-concept for the application of a penalized linear model to identify metalloprotein superfamilies on a large-scale. The application employed here, screening for iron-sulfur proteomes, provides new candidates for further biochemical and structural analysis as well as new resources for an extensive exploration of iron-sulfuromes in the microbial world.