Improving protein secondary structure prediction using a simple k-mer model.

Improving protein secondary structure prediction using a simple k-mer model.
复制标题

DOI:
10.1093/bioinformatics/btq020
复制
发表时间:
2010-03-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Gough J
Gough J
中科院分区:
其他
文献类型:
--
作者:
Madera M;Calmus R;Thiltgen G;Karplus K;Gough J

文献摘要

参考文献

被引文献

相似文献

动机:蛋白质序列分析的一些一阶方法固有地将每个位置视为独立的。我们开发了一个通用的框架,介绍更长的范围内的相互作用。然后,我们通过将其应用于二级结构预测来展示我们的方法的能力;在独立性假设下,由现有方法产生的序列可以产生不像蛋白质的特征,一个极端的例子是长度为1的螺旋。我们的目标是使最先进的方法的预测更加现实,而不损失其他措施的性能。结果:我们的框架更长的范围内的相互作用被描述为一个k-mer顺序模型。我们成功地将我们的模型应用于二级结构预测的特定问题,作为现有方法之上的附加层。我们实现了使预测更加真实和蛋白质化的目标,并且显着地提高了整体性能。我们将片段OVerlap(SOV)得分提高了1.8%,但更重要的是,我们从根本上提高了真实的序列的概率,从每个残基的平均预测值0.271提高到0.385。至关重要的是,这种改进是在没有额外信息的情况下获得的。可用性:http://supfam.cs.bris.ac.uk/kmer联系人:gough@cs.bris.ac.uk
Motivation: Some first order methods for protein sequence analysis inherently treat each position as independent. We develop a general framework for introducing longer range interactions. We then demonstrate the power of our approach by applying it to secondary structure prediction; under the independence assumption, sequences produced by existing methods can produce features that are not protein like, an extreme example being a helix of length 1. Our goal was to make the predictions from state of the art methods more realistic, without loss of performance by other measures. Results: Our framework for longer range interactions is described as a k-mer order model. We succeeded in applying our model to the specific problem of secondary structure prediction, to be used as an additional layer on top of existing methods. We achieved our goal of making the predictions more realistic and protein like, and remarkably this also improved the overall performance. We improve the Segment OVerlap (SOV) score by 1.8%, but more importantly we radically improve the probability of the real sequence given a prediction from an average of 0.271 per residue to 0.385. Crucially, this improvement is obtained using no additional information. Availability: http://supfam.cs.bris.ac.uk/kmer Contact: gough@cs.bris.ac.uk
DOI: 10.1093/bioinformatics/bti203
发表时间: 2005-04-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Pollastri, G;McLysaght, A
通讯作者: McLysaght, A
DOI: 10.1110/ps.9.6.1162
发表时间: 2000-06-01
期刊: PROTEIN SCIENCE
影响因子: 8
作者:
Ouali, M;King, RD
通讯作者: King, RD
DOI: 10.1093/bioinformatics/bth370
发表时间: 2004-11-22
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Liu, Y;Carbonell, J;Gopalakrishnan, V
通讯作者: Gopalakrishnan, V
DOI: 10.1006/jmbi.1993.1413
发表时间: 1993-07-20
影响因子: 5.6
作者:
ROST, B;SANDER, C
通讯作者: SANDER, C
DOI: 10.1093/bioinformatics/14.10.892
发表时间: 1998-01-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Cuff, JA;Clamp, ME;Barton, GJ
通讯作者: Barton, GJ