Comparison of five methods for finding conserved sequences in multiple alignments of gene regulatory regions

Comparison of five methods for finding conserved sequences in multiple alignments of gene regulatory regions
复制标题

DOI:
10.1093/nar/27.19.3899
复制
发表时间:
1999-10-01
影响因子:
14.9
通讯作者:
Hardison, R
Hardison, R
中科院分区:
生物学2区
文献类型:
--
作者:
Stojanovic, N;Florea, L;Hardison, R

文献摘要

被引文献

相似文献

DNA 或蛋白质序列中的保守片段是功能元件的有力候选者,因此需要开发和比较适当的计算方法。我们描述了五种方法和计算机程序,用于在先前计算的多重比对(主要针对 DNA 序列)中寻找高度保守的片段。其中两种方法已经得到普遍使用;这些基于良好的列一致性和高信息内容。另外三种方法找到具有最小进化变化的块、与已知中心序列每行至多 k 个位置不同的块以及与先验未知的中心序列每行至多 k 个位置不同的块。后两种方法中的中心序列是一种对 DNA 序列中已知或未知蛋白质的潜在结合位点进行建模的方法。通过分析哺乳动物β-珠蛋白基因簇中三个广泛分析的调控区和细菌阿拉伯糖操纵子的控制区来评估每种方法的功效。尽管所有五种方法具有完全不同的理论基础,但当调整其参数以最接近实验数据时,它们在这些数据集上产生相当相似的结果。基于信息内容的方法的最佳参数对于β-珠蛋白基因簇的不同调控区变化不大,因此可以外推到许多其他调控区。基于每行最大允许失配的程序具有简单的参数,其值可以先验地选择,因此当无法针对已知功能位点进行校准时,它们可能比其他方法更有用。
Conserved segments in DNA or protein sequences are strong candidates for functional elements and thus appropriate methods for computing them need to be developed and compared, We describe five methods and computer programs for finding highly conserved blocks within previously computed multiple alignments, primarily for DNA sequences. Two of the methods are already in common use; these are based on good column agreement and high information content, Three additional methods find blocks with minimal evolutionary change, blocks that differ in at most k positions per row from a known center sequence and blocks that differ in at most: k positions per row from a center sequence that is unknown a priori. The center sequence in the latter two methods is a way to model potential binding sites for known or unknown proteins in DNA sequences. The efficacy of each method was evaluated by analysis of three extensively analyzed regulatory regions in mammalian beta-globin gene clusters and the control region of bacterial arabinose operons, Although all five methods have quite different theoretical underpinnings, they produce rather similar results on these data sets when their parameters are adjusted to best approximate the experimental data, The optimal parameters for the method based on information content varied little for different regulatory regions of the beta-globin gene cluster and hence may be extrapolated to many other regulatory regions. The programs based on maximum allowed mismatches per row have simple parameters whose values can be chosen a priori and thus they may be more useful than the other methods when calibration against known functional sites is not available.