Transmembrane helix prediction using amino acid property features and latent semantic analysis.

Transmembrane helix prediction using amino acid property features and latent semantic analysis.
复制标题

DOI:
10.1186/1471-2105-9-s1-s4
复制
发表时间:
2008
期刊:
影响因子:
3
通讯作者:
Klein-Seetharaman J
Klein-Seetharaman J
中科院分区:
生物学4区
文献类型:
--
作者:
Ganapathiraju M;Balakrishnan N;Reddy R;Klein-Seetharaman J

文献摘要

被引文献

相似文献

用统计学方法预测跨膜(TM)螺旋缺乏足够的训练数据。目前最好的方法在模型中使用数百甚至数千个自由参数,这些参数经过调整以适应可用于训练的少量数据。此外,它们通常限于普遍接受的拓扑结构“细胞质-跨膜-细胞外”,并且不能适应不符合该拓扑结构的膜蛋白。最近的通道蛋白的晶体结构揭示了新的架构,表明上述拓扑结构可能不像以前认为的那样普遍。因此,有必要的方法,可以更好地预测TM螺旋,甚至在新的拓扑结构和家庭。在这里,我们描述了一种新的方法“TMpro”,以高精度预测TM螺旋。为了避免过度拟合现有的拓扑结构,我们已经折叠细胞质和细胞外标签到一个单一的状态,非TM。TMpro是一个二元分类器,它使用多种氨基酸性质(电荷、极性、芳香性、大小和电子性质)作为特征来预测TM或非TM。通过应用用于文本文档潜在语义分析的框架从序列信息中提取特征,并将其输入到学习TM和非TM片段之间区别的神经网络。该模型仅使用25个自由参数。在基准分析中,与不需要已知蛋白质进化谱的最佳方法相比,TMpro实现了95%的片段F分数,对应于50%的错误率降低。当应用于更新和更大的高分辨率数据集PDBTM和MPtopo时,性能也得到了提高。TMpro预测膜蛋白与不寻常的或有争议的TM结构(K+通道,水通道蛋白和HIV包膜糖蛋白)进行了讨论。TMpro在TM片段建模中使用非常少的自由参数,而不是在最先进的膜预测方法中使用非常大量的自由参数,但实现了非常高的片段精度。考虑到高分辨率跨膜信息仅可用于非常少的蛋白质,这是非常有利的。因此,预计TMpro的最大影响是在预测具有新拓扑结构的蛋白质中的TM片段。在此基础上,提出了一种新的蛋白质序列特征提取方法,即潜在语义分析模型。这种方法在当前背景下的成功表明,它可以在其他基于序列的分析问题中找到潜在的应用。 和
Prediction of transmembrane (TM) helices by statistical methods suffers from lack of sufficient training data. Current best methods use hundreds or even thousands of free parameters in their models which are tuned to fit the little data available for training. Further, they are often restricted to the generally accepted topology "cytoplasmic-transmembrane-extracellular" and cannot adapt to membrane proteins that do not conform to this topology. Recent crystal structures of channel proteins have revealed novel architectures showing that the above topology may not be as universal as previously believed. Thus, there is a need for methods that can better predict TM helices even in novel topologies and families. Here, we describe a new method "TMpro" to predict TM helices with high accuracy. To avoid overfitting to existing topologies, we have collapsed cytoplasmic and extracellular labels to a single state, non-TM. TMpro is a binary classifier which predicts TM or non-TM using multiple amino acid properties (charge, polarity, aromaticity, size and electronic properties) as features. The features are extracted from sequence information by applying the framework used for latent semantic analysis of text documents and are input to neural networks that learn the distinction between TM and non-TM segments. The model uses only 25 free parameters. In benchmark analysis TMpro achieves 95% segment F-score corresponding to 50% reduction in error rate compared to the best methods not requiring an evolutionary profile of a protein to be known. Performance is also improved when applied to more recent and larger high resolution datasets PDBTM and MPtopo. TMpro predictions in membrane proteins with unusual or disputed TM structure (K+ channel, aquaporin and HIV envelope glycoprotein) are discussed. TMpro uses very few free parameters in modeling TM segments as opposed to the very large number of free parameters used in state-of-the-art membrane prediction methods, yet achieves very high segment accuracies. This is highly advantageous considering that high resolution transmembrane information is available only for very few proteins. The greatest impact of TMpro is therefore expected in the prediction of TM segments in proteins with novel topologies. Further, the paper introduces a novel method of extracting features from protein sequence, namely that of latent semantic analysis model. The success of this approach in the current context suggests that it can find potential applications in other sequence-based analysis problems. and