Alignment-Independent Techniques for Protein Classification

Alignment-Independent Techniques for Protein Classification
复制标题

DOI:
10.2174/157016408786733770
复制
发表时间:
2008-12-01
期刊:
影响因子:
0.8
通讯作者:
Flower, Darren R.
Flower, Darren R.
中科院分区:
生物学4区
文献类型:
--
作者:
Davies, Matthew N.;Secker, Andrew;Flower, Darren R.

文献摘要

被引文献

相似文献

从氨基酸序列预测蛋白质的结构和功能是生物信息学的核心目标。大多数生物信息学分析使用序列比对作为测量相似性的基础。然而,越来越多的证据表明,许多蛋白质家族对这种简单的比较方法有抵抗力。越来越多地,机器学习技术和蛋白质序列的抽象表示的组合被用于基于其物理化学性质的相似性而不是评分序列比对来对蛋白质进行分类。这在显示出更大的结构保守性但似乎缺乏保守序列的蛋白质家族中特别有效。在这里,我们描述了固有的局限性依赖于蛋白质分类的方法,并提出了“无干扰”的表示作为一个可行的和现实的替代方案,以解决生物信息学中的复杂问题。
Predicting protein structure and function from amino acid sequences is a central aim of bioinformatics. Most bioinformatics analyses use sequence alignment as the basis by which to measure similarity. However, there is increasing evidence that many protein families are resistant to this straightforward method of comparison. Increasingly, a combination of machine-learning techniques and abstract representations of protein sequences is being used to classify proteins based upon the similarity of their physico-chemical properties rather than scoring sequence alignments. This is particularly effective in protein families that show greater structural conservation but appear to lack conserved sequences. Here we describe the inherent limitations of the alignment-dependent approaches to protein classification and present 'alignment-free' representations as a viable and realistic alternative to solve complex problems within bioinformatics.