Analyses of deep mammalian sequence alignments and constraint predictions for 1% of the human genome

Analyses of deep mammalian sequence alignments and constraint predictions for 1% of the human genome
复制标题

DOI:
10.1101/gr.6034307
复制
发表时间:
2007-06-01
期刊:
影响因子:
7
通讯作者:
Sidow, Arend
Sidow, Arend
中科院分区:
生物学1区
文献类型:
--
作者:
Margulies, Elliott H.;Cooper, Gregory M.;Sidow, Arend

文献摘要

被引文献

相似文献

正在进行的ENCODE项目的一个关键组成部分是对最初目标的1%的人类基因组进行严格的比较序列分析。在这里,我们提出了所有ENCODE目标的23种哺乳动物的正向序列生成,比对和进化约束分析。使用四种不同的方法产生比对;这些方法的比较揭示了大规模的一致性,但在小基因组重排、灵敏度(序列覆盖率)和特异性(比对准确度)方面存在实质性差异。我们描述了定量和定性的权衡与对齐方法的选择和水平的技术错误,需要占在应用程序中,需要多序列比对。使用生成的比对,我们使用三种不同的方法确定了约束区域。虽然不同的约束检测方法是在一般协议,有重要的差异有关的基本路线和具体的算法。然而,通过整合比对和约束检测方法的结果,我们产生了基于多个独立测量的约束注释。这些注释的分析表明,实验注释的功能元件的大多数类富集的约束序列,然而,每个类的大部分(蛋白质编码序列除外)不重叠的约束区域。后一种元件可能不受一级序列约束,可能不受所有哺乳动物的约束,或者可能具有消耗性分子功能。相反,40%的受约束序列不与实验鉴定的任何功能元件重叠。总之,这些发现证明并量化了有多少基因组功能元件等待基本的分子表征。
A key component of the ongoing ENCODE project involves rigorous comparative sequence analyses for the initially targeted 1% of the human genome. Here, we present orthologous sequence generation, alignment, and evolutionary constraint analyses of 23 mammalian species for all ENCODE targets. Alignments were generated using four different methods; comparisons of these methods reveal large-scale consistency but substantial differences in terms of small genomic rearrangements, sensitivity ( sequence coverage), and specificity ( alignment accuracy). We describe the quantitative and qualitative trade-offs concomitant with alignment method choice and the levels of technical error that need to be accounted for in applications that require multisequence alignments. Using the generated alignments, we identified constrained regions using three different methods. While the different constraint-detecting methods are in general agreement, there are important discrepancies relating to both the underlying alignments and the specific algorithms. However, by integrating the results across the alignments and constraint-detecting methods, we produced constraint annotations that were found to be robust based on multiple independent measures. Analyses of these annotations illustrate that most classes of experimentally annotated functional elements are enriched for constrained sequences; however, large portions of each class ( with the exception of protein-coding sequences) do not overlap constrained regions. The latter elements might not be under primary sequence constraint, might not be constrained across all mammals, or might have expendable molecular functions. Conversely, 40% of the constrained sequences do not overlap any of the functional elements that have been experimentally identified. Together, these findings demonstrate and quantify how many genomic functional elements await basic molecular characterization.