An efficient, versatile and scalable pattern growth approach to mine frequent patterns in unaligned protein sequences

An efficient, versatile and scalable pattern growth approach to mine frequent patterns in unaligned protein sequences
复制标题

DOI:
10.1093/bioinformatics/btl665
复制
发表时间:
2007-03-15
期刊:
影响因子:
5.8
通讯作者:
Ijzerman, Adriaan P.
Ijzerman, Adriaan P.
中科院分区:
生物学3区
文献类型:
--
作者:
Ye, Kai;Kosters, Walter A.;Ijzerman, Adriaan P.

文献摘要

被引文献

相似文献

动机:蛋白质序列中的模式发现通常基于多序列比对(MSA)。该过程可能是计算密集型的,并且通常需要手动调整,这对于一组偏离序列可能特别困难。相比之下,两种算法PRATT 2(http//www.ebi.ac.uk/pratt/)和TEIRESIAS(http://cbcsrv.watson.ibm.com/)用于直接从未比对的生物序列中识别频繁模式,而不尝试将它们进行比对。在这里,我们提出了一个新的算法,更高效,更功能比PRATT 2和TEIRESIAS,并讨论了它的一些应用G蛋白偶联受体,一个蛋白质家族的重要drug targets.Results:在这项研究中,我们设计并实现了六个算法,挖掘三种不同的模式类型,从一个或两个数据集使用模式增长的方法。我们比较了我们的方法,PRATT 2和TEIRESIAS的效率,完整性和模式类型的多样性。与PRATT 2相比,我们的方法更快,能够处理大型数据集,并能够识别所谓的III型模式。我们的方法在发现所谓的I型模式方面与TEIRESIAS相当,但具有额外的功能,例如挖掘所谓的II型和III型模式,并找到两个数据集之间的区别模式。
Motivation: Pattern discovery in protein sequences is often based on multiple sequence alignments (MSA). The procedure can be computationally intensive and often requires manual adjustment, which may be particularly difficult for a set of deviating sequences. In contrast, two algorithms, PRATT2 (http//www.ebi.ac.uk/pratt/) and TEIRESIAS (http://cbcsrv.watson.ibm.com/) are used to directly identify frequent patterns from unaligned biological sequences without an attempt to align them. Here we propose a new algorithm with more efficiency and more functionality than both PRATT2 and TEIRESIAS, and discuss some of its applications to G protein-coupled receptors, a protein family of important drug targets.Results: In this study, we designed and implemented six algorithms to mine three different pattern types from either one or two datasets using a pattern growth approach. We compared our approach to PRATT2 and TEIRESIAS in efficiency, completeness and the diversity of pattern types. Compared to PRATT2, our approach is faster, capable of processing large datasets and able to identify the so-called type III patterns. Our approach is comparable to TEIRESIAS in the discovery of the so-called type I patterns but has additional functionality such as mining the so-called type II and type III patterns and finding discriminating patterns between two datasets.