The relationship between classification of multi-domain proteins using an alignment-free approach and their functions: a case study with immunoglobulins

The relationship between classification of multi-domain proteins using an alignment-free approach and their functions: a case study with immunoglobulins
复制标题

DOI:
10.1039/c3mb70443b
复制
发表时间:
2014-01-01
影响因子:
--
通讯作者:
Srinivasan, Narayanaswamy
Srinivasan, Narayanaswamy
中科院分区:
生物3区
文献类型:
--
作者:
Bhaskara, Ramachandra M.;Mehrotra, Prachi;Srinivasan, Narayanaswamy

文献摘要

被引文献

相似文献

建立多结构域蛋白质序列之间的功能关系是一项重要的任务。传统上,描述蛋白质的功能分配和关系需要结构域分配作为先决条件。该过程对对齐质量和域定义敏感。在多结构域蛋白质中,由于多种原因,比对的质量较差。我们报告的对应关系的分类蛋白质表示为全长基因产物和它们的功能。我们的方法从根本上不同于传统的方法,在不执行域的级别上的分类。我们的方法是基于在氨基酸序列水平上的对齐自由本地匹配分数(LMS)计算,然后进行层次聚类。由于全长蛋白质序列分类没有金标准,我们采用基因本体和基于域架构的相似性度量来评估我们的分类。使用LMS获得的最终聚类显示出高的功能和域架构相似性。在结构域和全长蛋白质上将当前方法与基于比对的方法进行比较,显示出LMS分数的优越性。使用这种方法,我们已经重新创建了不同的蛋白激酶亚家族之间的客观关系,也分类的免疫球蛋白含有的蛋白质亚家族的定义目前不存在。这种方法可以应用于任何一组蛋白质序列,因此将有助于分析大量的全长蛋白质序列。
Establishing functional relationships between multi-domain protein sequences is a non-trivial task. Traditionally, delineating functional assignment and relationships of proteins requires domain assignments as a prerequisite. This process is sensitive to alignment quality and domain definitions. In multi-domain proteins due to multiple reasons, the quality of alignments is poor. We report the correspondence between the classification of proteins represented as full-length gene products and their functions. Our approach differs fundamentally from traditional methods in not performing the classification at the level of domains. Our method is based on an alignment free local matching scores (LMS) computation at the amino-acid sequence level followed by hierarchical clustering. As there are no gold standards for full-length protein sequence classification, we resorted to Gene Ontology and domain-architecture based similarity measures to assess our classification. The final clusters obtained using LMS show high functional and domain architectural similarities. Comparison of the current method with alignment based approaches at both domain and full-length protein showed superiority of the LMS scores. Using this method we have recreated objective relationships among different protein kinase sub-families and also classified immunoglobulin containing proteins where sub-family definitions do not exist currently. This method can be applied to any set of protein sequences and hence will be instrumental in analysis of large numbers of full-length protein sequences.