Finding functional sequence elements by multiple local alignment

Finding functional sequence elements by multiple local alignment
复制标题

DOI:
10.1093/nar/gkh169
复制
发表时间:
2004-01-01
影响因子:
14.9
通讯作者:
Weng, ZP
Weng, ZP
中科院分区:
生物学2区
文献类型:
--
作者:
Frith, MC;Hansen, U;Weng, ZP

文献摘要

被引文献

相似文献

检测和比对生物序列局部相似区域的算法有可能发现各种各样的功能基序。这里针对这一经典但未解决的问题提出了两项理论贡献:一种自动确定比对基序宽度的方法;一种计算比对统计显著性的技术,即评估比对是否比在随机、不相关序列中偶然出现的比对更强。在探索标准吉布斯采样技术的变体以优化比对时,我们发现模拟退火方法更有效。最后,我们通过将算法应用于越来越困难的测试案例进行失败测试,并分析最终失败的方式和原因。转录因子结合基序的检测受到基序内在微妙性的限制,而不是比对优化过程的不足。
Algorithms that detect and align locally similar regions of biological sequences have the potential to discover a wide variety of functional motifs. Two theoretical contributions to this classic but unsolved problem are presented here: a method to determine the width of the aligned motif automatically; and a technique for calculating the statistical significance of alignments, i.e. an assessment of whether the alignments are stronger than those that would be expected to occur by chance among random, unrelated sequences. Upon exploring variants of the standard Gibbs sampling technique to optimize the alignment, we discovered that simulated annealing approaches perform more efficiently. Finally, we conduct failure tests by applying the algorithm to increasingly difficult test cases, and analyze the manner of and reasons for eventual failure. Detection of transcription factor-binding motifs is limited by the motifs' intrinsic subtlety rather than by inadequacy of the alignment optimization procedure.