Computational screening of conserved genomic DNA in search of functional noncoding elements

Computational screening of conserved genomic DNA in search of functional noncoding elements
复制标题

DOI:
10.1038/nmeth0705-535
复制
发表时间:
2005-07-01
期刊:
影响因子:
48
通讯作者:
Haussler, D
Haussler, D
中科院分区:
生物学1区
文献类型:
--
作者:
Bejerano, G;Siepel, AC;Haussler, D

文献摘要

被引文献

相似文献

小鼠基因组的测序首次允许对我们自己基因组内序列保守程度的大规模估计。特别是,它表明,在哺乳动物中,保守的基因组DNA至少是蛋白质编码DNA的两倍(1)。保守的非编码区的丰富性甚至适用于哺乳动物保护尺度最顶端的所谓超保守元件(2),以及在长分化脊椎动物(如人类和鱼类)之间保守的序列(3,4)。观察到的大部分保守性似乎是纯化选择的结果,这表明存在大量未表征的功能元件和家族(5),包括转录和转录后调控元件、染色质结构相关区域、非编码RNA,以及可能完全是新的功能元件类别(6)。此外,最近的比较测序工作揭示了其他后生动物中类似的丰富的未表征的保守非编码序列(7)。我们在这里概述了如何从广泛的模式生物中获得保守区域的集合。然后,我们描述了如何分析这些区域的属性,过滤掉不需要的,如已知的和预测的编码区域,并排名为进一步的计算和功能分析的其余部分。该方法大量使用UCSC基因组浏览器数据库(8)和一套相关的网络访问工具。该协议分为四个主要步骤:定义感兴趣的基因组区域,基于用户的起点(感兴趣的基因,两个遗传标记之间的区域,和其他区域);基于隐马尔可夫模型选择该区域内的跨物种保守元件的子集,所述隐马尔可夫模型定义并评分用于保守的基因组间隔;映射区间组的不同性质(例如转录物重叠和物种覆盖范围);以及最后,基于感兴趣的功能类别的特征谱,对该组进行排序以用于进一步分析。相同的协议可用于在UCSC基因组浏览器中可用的生命树的所有分支中搜索不同功能类别的元件,包括脊椎动物、昆虫、线虫和酵母。它还可以很容易地将用户可以访问的自定义类型的信息,并允许容易地更换部分协议,因为我们的理解之间的关系的功能和序列的保护,以及不同的功能类,提高。作为一个例子,我们提出了脊椎动物增强子序列的信息学概况,并讨论了这种方法导致发现几个功能增强子的情况。一个附带的方案描述了一种互补的方法,通过聚类对应于已知转录因子结合位点的序列基序来鉴定复杂基因组组装体中的顺式调控DNA区域(9)。
The sequencing of the mouse genome allowed, for the first time, the large-scale estimation of the extent of sequence conservation within our own genome. In particular, it suggested that in mammals there is at Least twice as much conserved genomic DNA as there is protein coding DNA(1). The abundance of conserved noncoding regions holds even for so-called ultraconserved elements at the very tip of the mammalian conservation scale(2), as well as for sequences conserved between long-diverged vertebrates such as human and fish(3,4). Much of the observed conservation appears to be the result of purifying selection, suggesting a wealth of uncharacterized functional elements and families(5), including transcriptional and post-transcriptional regulatory elements, chromatin structure-associated regions, noncoding RNAs, and perhaps altogether novel classes of functional etements(6). Moreover, recent comparative sequencing efforts have revealed similarly rich sets of uncharacterized conserved noncoding sequences in other metazoans(7). We outline here how to obtain sets of conserved regions from a wide range of model organisms. We then describe how to analyze the properties of these regions, filter out undesired ones, such as known and predicted coding regions, and rank the remainder for further computational and functional analysis. The method makes heavy use of the UCSC Genome Browser Database(8) and a suite of related web-accessible tools. The protocol is divided into four major steps: defining the genomic region of interest, based on the user's starting point (gene of interest, a region between two genetic markers, and other regions); selecting a subset of cross-species conserved elements within this region, based on a hidden Markov model that defines and scores genomic intervals for conservation; mapping the different properties of the interval set (such as transcript overlap and species coverage extent); and, finally, ranking the set for further analysis, based on a characteristic profile of the functional class of interest. The same protocol may be used to search for different functional classes of elements in all branches of the tree of life available in the UCSC Genome Browser, including vertebrate, insect, nematode and yeast. It can also easily incorporate custom types of information that the user has access to and allows for easy replacement of parts of the protocol, as our understanding of the relationship between function and sequence conservation, and of the different functional classes, improves. As an example, we present an informatic profile of vertebrate enhancer sequences and discuss a case for which such a method has led to the discovery of several functional enhancers. An accompanying protocol describes a complementary approach to identification of cis-regulatory DNA regions in complex genome assemblies by clustering of sequence motifs corresponding to known transcription factor binding sites(9).