Identification and characterization of multi-species conserved sequences

Identification and characterization of multi-species conserved sequences
复制标题

DOI:
10.1101/gr.1602203
复制
发表时间:
2003-12-01
期刊:
影响因子:
7
通讯作者:
Green, ED
Green, ED
中科院分区:
生物学1区
文献类型:
--
作者:
Margulies, EH;Blanchette, M;Green, ED

文献摘要

被引文献

相似文献

比较序列分析已成为阐明基因组功能的重要组成部分。来自多种脊椎动物的基因组序列的日益增加的可用性产生了对可以以稳健的方式检测高度保守区域的计算方法的需求。为此,我们正在开发方法来识别在多个物种中保守的序列;我们称之为“多物种保守序列”(或MCS)。在这里,我们报告了两种策略的MCS识别,证明了他们的能力,几乎所有已知的积极保守的序列(具体而言,编码序列),但很少中性进化序列(具体而言,祖先重复)。重要的是,我们发现MCSs中相当一部分碱基(约70%)位于非编码区;因此,在多个脊椎动物物种中保守的大部分序列没有已知的功能。这些MCS的初步表征揭示了对应于转录因子结合位点簇、非编码RNA转录物和其他候选功能元件的序列。最后,检测MCSs的能力代表了评估物种序列对识别感兴趣的基因组区域的相对贡献的有价值的度量,并且我们的结果表明,目前可用的基因组序列不足以全面识别人类基因组中的MCSs。
Comparative sequence analysis has become an essential component of studies aiming to elucidate genome function. The increasing availability of genomic sequences from multiple vertebrates is creating the need for computational methods that can detect highly conserved regions in a robust fashion. Towards that end, we are developing approaches for identifying sequences that are conserved across multiple species; we call these "Multi-species Conserved Sequences" (or MCSs). Here we report two strategies for MCS identification, demonstrating their ability to detect virtually all known actively conserved sequences (specifically, coding sequences) but very little neutrally evolving sequence (specifically, ancestral repeats). Importantly, we find that a substantial fraction of the bases within MCSs (similar to70%) resides within non-coding regions; thus, the majority of sequences conserved across multiple vertebrate species has no known function. Initial characterization of these MCSs has revealed sequences that correspond to clusters of transcription factor-binding sites, non-coding RNA transcripts, and other candidate functional elements. Finally, the ability to detect MCSs represents a valuable metric for assessing the relative contribution of a species' sequence to identifying genomic regions of interest, and our results indicate that the currently available genome sequences are insufficient for the comprehensive identification of MCSs in the human genome.