Repetitive elements may comprise over two-thirds of the human genome.

Repetitive elements may comprise over two-thirds of the human genome.
复制标题

DOI:
10.1371/journal.pgen.1002384
复制
发表时间:
2011-12
期刊:
影响因子:
4.5
通讯作者:
Pollock DD
Pollock DD
中科院分区:
生物学2区
文献类型:
--
作者:
de Koning AP;Gu W;Castoe TA;Batzer MA;Pollock DD

文献摘要

参考文献

被引文献

相似文献

转座元件(TEs)在真核生物基因组中通常通过与共有元件序列比对来鉴定。使用这种方法,人类基因组中约一半先前已被鉴定为转座元件和低复杂度重复序列。我们最近开发了一种高度灵敏的全新替代策略——P - 云,它转而搜索在序列空间上相关的高丰度寡核苷酸簇(寡核苷酸“云”)。我们在此表明,P - 云预测人类基因组中还有>840 Mbp的重复序列,因此表明人类基因组的66% - 69%是重复的或由重复序列衍生而来。为了研究这种显著差异,我们对P - 云以及一种常用的常规方法——RepeatMasker(RM)检测高度丰富的人类Alu和MIR短散在核元件(SINEs)不同大小片段的能力进行了详细分析。与P - 云不同,RM对于即使是中等长度的片段也可能具有令人惊讶的低灵敏度,而P - 云对于小至约25 bp的片段大小仍具有良好的灵敏度。尽管短片段具有较高的假阳性内在概率,但我们进行了一种反映这一事实的概率注释。我们进一步开发了“元件特异性”P - 云(ESPs)来鉴定新的Alu和MIR SINE元件,并且使用它我们鉴定出约100 Mb先前未注释的人类元件。对新MIR序列的ESP估计与基于RM对RM所遗漏数量的预测非常吻合。这些结果强调了需要结合概率性的基因组注释方法,并表明人类基因组包含的重复序列比先前认为的要多得多。 我们的研究关注关于人类基因组序列的一个基本问题:它是由什么组成的?目前,大约50%的基因组序列功能或起源未知,有时被称为人类基因组的“暗物质”。我们在此证明,这片未知领域中大约一半实际上是由重复的或由重复序列衍生而来的序列组成,这些序列很可能以转座元件为主。这些序列差异过大或退化,以至于难以通过与已知转座元件共有序列比对来检测,但可以使用我们在此提出并评估的新型全新搜索方法来检测。因此,检测重复序列的标准方法可能遗漏了大量转座元件片段。在一种情况(MIR元件)下,我们预测基因组中可能存在的一半序列未被检测到。在全基因组范围内,在控制假阳性之后,我们推断人类基因组序列的绝大多数(>66% - 69%)是由重复的和由重复序列衍生而来的DNA元件组成。这一估计与先前的估计形成鲜明对比,并表明转座元件在塑造我们基因组的历史和内容方面所起的作用比先前认为的要大得多。
Transposable elements (TEs) are conventionally identified in eukaryotic genomes by alignment to consensus element sequences. Using this approach, about half of the human genome has been previously identified as TEs and low-complexity repeats. We recently developed a highly sensitive alternative de novo strategy, P-clouds, that instead searches for clusters of high-abundance oligonucleotides that are related in sequence space (oligo “clouds”). We show here that P-clouds predicts >840 Mbp of additional repetitive sequences in the human genome, thus suggesting that 66%–69% of the human genome is repetitive or repeat-derived. To investigate this remarkable difference, we conducted detailed analyses of the ability of both P-clouds and a commonly used conventional approach, RepeatMasker (RM), to detect different sized fragments of the highly abundant human Alu and MIR SINEs. RM can have surprisingly low sensitivity for even moderately long fragments, in contrast to P-clouds, which has good sensitivity down to small fragment sizes (∼25 bp). Although short fragments have a high intrinsic probability of being false positives, we performed a probabilistic annotation that reflects this fact. We further developed “element-specific” P-clouds (ESPs) to identify novel Alu and MIR SINE elements, and using it we identified ∼100 Mb of previously unannotated human elements. ESP estimates of new MIR sequences are in good agreement with RM-based predictions of the amount that RM missed. These results highlight the need for combined, probabilistic genome annotation approaches and suggest that the human genome consists of substantially more repetitive sequence than previously believed. Our study is concerned with a fundamental question about the human genome sequence: what is it made of? At present, approximately 50% of the genome sequence has unknown function or origin and is sometimes referred to as the “dark matter” of the human genome. We demonstrate here that approximately half of this uncharted territory is in fact comprised of repetitive or repeat-derived sequences, which are most likely dominated by transposable elements. These sequences are too diverged or degraded to be easily detected by alignment to known transposable element consensus sequences, but can be detected using novel de novo search methods that we present and evaluate here. Standard methods for detecting repetitive sequence are therefore probably missing large numbers of transposable element fragments. In one case (MIR elements), we predict that half of the sequence that is likely present in the genome has gone undetected. Genome-wide, we infer that a large majority of the human genome sequence (>66%–69%) is comprised of repetitive and repeat-derived DNA elements, after controlling for false positives. This estimate stands in stark contrast with previous estimates and suggests that transposable elements have played a much larger role in shaping the history and content of our genome than previously believed.
DOI: 10.1371/journal.pcbi.0010043
发表时间: 2005-09
影响因子: 4.3
作者:
Li R;Ye J;Li S;Wang J;Han Y;Ye C;Wang J;Yang H;Yu J;Wong GK;Wang J
通讯作者: Wang J
DOI: 10.1126/science.1086432
发表时间: 2003-09-26
期刊: SCIENCE
影响因子: 56.9
作者:
Kirkness, EF;Bafna, V;Venter, JC
通讯作者: Venter, JC
DOI: 10.1093/bioinformatics/bti1003
发表时间: 2005-06-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Edgar, RC;Myers, EW
通讯作者: Myers, EW
DOI: 10.1101/gr.6725608
发表时间: 2008-02-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Lunter, Gerton;Rocco, Andrea;Hein, Jotun
通讯作者: Hein, Jotun
DOI: 10.1093/gbe/evr043
发表时间: 2011
影响因子: 3.3
作者:
Castoe TA;Hall KT;Guibotsy Mboulas ML;Gu W;de Koning AP;Fox SE;Poole AW;Vemulapalli V;Daza JM;Mockler T;Smith EN;Feschotte C;Pollock DD
通讯作者: Pollock DD