FastCluster: a graph theory based algorithm for removing redundant sequences

FastCluster: a graph theory based algorithm for removing redundant sequences
复制标题

FastCluster:一种基于图论的冗余序列去除算法

DOI:
10.4236/jbise.2009.28090
复制
发表时间:
2009-12
期刊:
J. Biomedical Science and Engineering
影响因子:
--
通讯作者:
董骝焕
董骝焕
中科院分区:
其他
文献类型:
--
作者:
董骝焕

文献摘要

参考文献

相似文献

在许多情况下,生物序列数据库包含冗余序列,这使得实现可靠的统计分析变得困难。从一个庞大的序列数据集中剔除冗余序列,找出所有真实的蛋白质家族及其代表,在生物信息学中是非常重要的。去除冗余蛋白质序列的问题可以建模为从图中寻找最大独立集,这在数学上是一个NP问题。本文以数学图论为基础,提出了一种新的程序--FastCluster.该算法对Hobohm和Sander的算法进行了改进,以生成无冗余的蛋白质序列集。FastCluster使用BLAST来确定两个序列之间的相似性,以获得更好的序列相似性。将该算法的性能与Hobohm和Sander的算法进行了比较,结果表明,该算法能产生一个合理的无冗余集合,相似度下限为0.0~1.0。该算法在生成更大的最大非冗余(独立)蛋白质集时更接近实际结果(图的最大独立集),这意味着所有的蛋白质家族都是聚类的。这使得快速聚类成为删除冗余蛋白质序列的宝贵工具。
In many cases, biological sequence databases contain redundant sequences that make it difficult to achieve reliable statistical analysis. Removing the redundant sequences to find all the real protein families and their representatives from a large sequences dataset is quite important in bioinformatics. The problem of removing redundant protein sequences can be modeled as finding the maximum independent set from a graph, which is a NP problem in Mathematics. This paper presents a novel program named FastCluster on the basis of mathematical graph theory. The algorithm makes an improvement to Hobohm and Sander’s algorithm to generate non-redundant protein sequence sets. FastCluster uses BLAST to determine the similarity between two sequences in order to get better sequence similarity. The algorithm’s performance is compared with Hobohm and Sander’s algorithm and it shows that Fast- Cluster can produce a reasonable non-redundant pro- tein set and have a similarity cut-off from 0.0 to 1.0. The proposed algorithm shows its superiority in generating a larger maximal non-redundant (independent) protein set which is closer to the real result (the maximum independent set of a graph) that means all the protein families are clustered. This makes Fast- Cluster a valuable tool for removing redundant protein sequences.
DOI: 10.1002/pro.5560030317
发表时间: 1994-03
期刊: Protein Science
影响因子: 8
作者:
U. Hobohm;C. Sander
通讯作者: U. Hobohm;C. Sander
DOI: --
发表时间: 2003
期刊: --
影响因子: --
作者:
Sampo Niskanen;P. Östergård
通讯作者: Sampo Niskanen;P. Östergård
DOI: 10.1002/pro.5560010313
发表时间: 1992-03
期刊: Protein Science
影响因子: 8
作者:
U. Hobohm;M. Scharf;R. Schneider;C. Sander
通讯作者: U. Hobohm;M. Scharf;R. Schneider;C. Sander
DOI: 10.1093/bioinformatics/17.3.282
发表时间: 2001-03-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Li, WZ;Jaroszewski, L;Godzik, A
通讯作者: Godzik, A
DOI: 10.1515/iupac.90.0375
发表时间: 2019-03
期刊: IUPAC Standards Online
影响因子: --
作者:
J. Labuda;R. Bowater;M. Fojta;G. Gauglitz;Z. Glatz;I. Hapala;J. Havliš;F. Kilár;Anikó Kilár;Lenka Malinovská;Heli M. M. Sirén-Heli-M.-M.-Sirén-2136896047;P. Skládal;F. Torta;M. Valachovič;M. Wimmerová;Z. Zdráhal;D. B. Hibbert
通讯作者: J. Labuda;R. Bowater;M. Fojta;G. Gauglitz;Z. Glatz;I. Hapala;J. Havliš;F. Kilár;Anikó Kilár;Lenka Malinovská;Heli M. M. Sirén-Heli-M.-M.-Sirén-2136896047;P. Skládal;F. Torta;M. Valachovič;M. Wimmerová;Z. Zdráhal;D. B. Hibbert