Into the heart of darkness: large-scale clustering of human non-coding DNA

Into the heart of darkness: large-scale clustering of human non-coding DNA
复制标题

DOI:
10.1093/bioinformatics/bth946
复制
发表时间:
2004-08-04
期刊:
影响因子:
5.8
通讯作者:
Blanchette, Mathieu
Blanchette, Mathieu
中科院分区:
生物学3区
文献类型:
--
作者:
Bejerano, Gill;Haussler, David;Blanchette, Mathieu

文献摘要

被引文献

相似文献

动机:目前认为,人类基因组包含的非编码功能区域大约是蛋白质编码基因的两倍,但我们对这些区域的了解非常有限。结果:我们研究了人类、小鼠和大鼠基因组中同步保守序列之间的交集,以及人类基因组本身的序列相似性,以寻找非蛋白编码元件的家族。为此,我们开发了一种图论聚类算法,类似于在阐明蛋白质序列家族关系中使用的非常成功的方法。该算法应用于一组高度过滤的大约70万个人类-啮齿动物进化保守区域,这些区域与任何已知的编码序列都不相似,占人类基因组的3.7%。从这些,我们得到了大约12000个非单例簇,密集的显著序列相似性。对基因组位置、转录证据和RNA二级结构的进一步分析表明,许多集群在一个或多个特征上具有显著的同质性。人类基因组中高度保守的非蛋白编码元件的这一子集因此包含丰富的类家族结构,值得深入分析。
Motivation: It is currently believed that the human genome contains about twice as much non-coding functional regions as it does protein-coding genes, yet our understanding of these regions is very limited.Results: We examine the intersection between syntenically conserved sequences in the human, mouse and rat genomes, and sequence similarities within the human genome itself, in search of families of non-protein-coding elements. For this purpose we develop a graph theoretic clustering algorithm, akin to the highly successful methods used in elucidating protein sequence family relationships.The algorithm is applied to a highly filtered set of about 700 000 human-rodent evolutionarily conserved regions, not resembling any known coding sequence, which encompasses 3.7% of the human genome. From these, we obtain roughly 12 000 non-singleton clusters, dense in significant sequence similarities. Further analysis of genomic location, evidence of transcription and RNA secondary structure reveals many clusters to be significantly homogeneous in one or more characteristics. This subset of the highly conserved non-protein-coding elements in the human genome thus contains rich family-like structures, which merit in-depth analysis.