Genomic multiple sequence alignments: refinement using a genetic algorithm.

Genomic multiple sequence alignments: refinement using a genetic algorithm.
复制标题

DOI:
10.1186/1471-2105-6-200
复制
发表时间:
2005-08-08
期刊:
影响因子:
3
通讯作者:
Lefkowitz, EJ
Lefkowitz, EJ
中科院分区:
生物学4区
文献类型:
--
作者:
Wang, CL;Lefkowitz, EJ

文献摘要

参考文献

被引文献

相似文献

基因组序列数据不能完全孤立地被理解。比较基因组学--比较不同物种的基因组序列的实践--在理解导致表型差异的物种之间的基因差异以及揭示进化关系的模式方面发挥着越来越重要的作用。比较基因组学的主要挑战之一是在两个或更多相关基因组序列之间产生高质量的比对。近年来,已经开发了许多工具来比对大的基因组序列。大多数利用启发式策略来识别一系列强序列相似性,然后将其用作锚点之间的区域比对。由此产生的比对是全局正确的,但在许多情况下局部是次优的。我们描述了一个新的程序GenAlignRefine,它通过使用遗传算法来改善局部比对区域,从而提高了全局多重比对的整体质量。低质量的区域被识别,使用T-Coffee程序重新排列,然后使用遗传算法进行细化。由于较好的CAFO(基于一致性的比对评估目标函数)得分通常反映较好的比对质量,因此该算法搜索产生较好的CAFIC得分的比对。为了改善遗传算法的固有速度,GenAlignRefiny被实现为一个并行的、基于集群的程序。我们通过在Linux集群上运行GenAlignRefining算法来测试它,以从模拟中提炼序列,以及提炼最初由Multi-Lagan比对的15个正痘病毒基因组序列的多重比对,这些序列的长度约为260,000个核苷酸。一个拥有40个处理器的Linux集群花了大约150分钟来优化正痘病毒比对的大约200个模糊(排列不佳)区域。总体序列同一性仅略有增加;但值得注意的是,通过消除缺口,总比对长度减少了约200个缺口区域,代表大约1300个缺口。我们在并行模式下实现了一种遗传算法来优化最初由各种比对工具生成的多个基因组序列比对。基准实验表明,改进算法在一段合理的时间内改善了基因组序列的比对。
Genomic sequence data cannot be fully appreciated in isolation. Comparative genomics – the practice of comparing genomic sequences from different species – plays an increasingly important role in understanding the genotypic differences between species that result in phenotypic differences as well as in revealing patterns of evolutionary relationships. One of the major challenges in comparative genomics is producing a high-quality alignment between two or more related genomic sequences. In recent years, a number of tools have been developed for aligning large genomic sequences. Most utilize heuristic strategies to identify a series of strong sequence similarities, which are then used as anchors to align the regions between the anchor points. The resulting alignment is globally correct, but in many cases is suboptimal locally. We describe a new program, GenAlignRefine, which improves the overall quality of global multiple alignments by using a genetic algorithm to improve local regions of alignment. Regions of low quality are identified, realigned using the program T-Coffee, and then refined using a genetic algorithm. Because a better COFFEE (Consistency based Objective Function For alignmEnt Evaluation) score generally reflects greater alignment quality, the algorithm searches for an alignment that yields a better COFFEE score. To improve the intrinsic slowness of the genetic algorithm, GenAlignRefine was implemented as a parallel, cluster-based program. We tested the GenAlignRefine algorithm by running it on a Linux cluster to refine sequences from a simulation, as well as refine a multiple alignment of 15 Orthopoxvirus genomic sequences approximately 260,000 nucleotides in length that initially had been aligned by Multi-LAGAN. It took approximately 150 minutes for a 40-processor Linux cluster to optimize some 200 fuzzy (poorly aligned) regions of the orthopoxvirus alignment. Overall sequence identity increased only slightly; but significantly, this occurred at the same time that the overall alignment length decreased – through the removal of gaps – by approximately 200 gapped regions representing roughly 1,300 gaps. We have implemented a genetic algorithm in parallel mode to optimize multiple genomic sequence alignments initially generated by various alignment tools. Benchmarking experiments showed that the refinement algorithm improved genomic sequence alignments within a reasonable period of time.
DOI: 10.1093/nar/24.8.1515
发表时间: 1996-04-15
影响因子: 14.9
作者:
Notredame, C;Higgins, DG
通讯作者: Higgins, DG
DOI: 10.1093/bioinformatics/14.2.157
发表时间: 1998-01-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Stoye, J;Evers, D;Meyer, F
通讯作者: Meyer, F
DOI: 10.1186/1471-2105-5-6
发表时间: 2004-01-21
期刊: BMC bioinformatics
影响因子: 3
作者:
Pollard DA;Bergman CM;Stoye J;Celniker SE;Eisen MB
通讯作者: Eisen MB
DOI: 10.1093/bioinformatics/15.3.211
发表时间: 1999-03-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Morgenstern, B
通讯作者: Morgenstern, B
DOI: 10.1093/nar/22.22.4673
发表时间: 1994-11-11
影响因子: 14.9
作者:
THOMPSON, JD;HIGGINS, DG;GIBSON, TJ
通讯作者: GIBSON, TJ