Repetitive elements may comprise over two-thirds of the human genome.
Repetitive elements may comprise over two-thirds of the human genome.
复制标题
DOI:
10.1371/journal.pgen.1002384
复制
发表时间:
2011-12
期刊:
影响因子:
4.5
通讯作者:
Pollock DD
中科院分区:
文献类型:
--
作者:
de Koning AP;Gu W;Castoe TA;Batzer MA;Pollock DD
Transposable elements (TEs) are conventionally identified in eukaryotic genomes by alignment to consensus element sequences. Using this approach, about half of the human genome has been previously identified as TEs and low-complexity repeats. We recently developed a highly sensitive alternative de novo strategy, P-clouds, that instead searches for clusters of high-abundance oligonucleotides that are related in sequence space (oligo “clouds”). We show here that P-clouds predicts >840 Mbp of additional repetitive sequences in the human genome, thus suggesting that 66%–69% of the human genome is repetitive or repeat-derived. To investigate this remarkable difference, we conducted detailed analyses of the ability of both P-clouds and a commonly used conventional approach, RepeatMasker (RM), to detect different sized fragments of the highly abundant human Alu and MIR SINEs. RM can have surprisingly low sensitivity for even moderately long fragments, in contrast to P-clouds, which has good sensitivity down to small fragment sizes (∼25 bp). Although short fragments have a high intrinsic probability of being false positives, we performed a probabilistic annotation that reflects this fact. We further developed “element-specific” P-clouds (ESPs) to identify novel Alu and MIR SINE elements, and using it we identified ∼100 Mb of previously unannotated human elements. ESP estimates of new MIR sequences are in good agreement with RM-based predictions of the amount that RM missed. These results highlight the need for combined, probabilistic genome annotation approaches and suggest that the human genome consists of substantially more repetitive sequence than previously believed. Our study is concerned with a fundamental question about the human genome sequence: what is it made of? At present, approximately 50% of the genome sequence has unknown function or origin and is sometimes referred to as the “dark matter” of the human genome. We demonstrate here that approximately half of this uncharted territory is in fact comprised of repetitive or repeat-derived sequences, which are most likely dominated by transposable elements. These sequences are too diverged or degraded to be easily detected by alignment to known transposable element consensus sequences, but can be detected using novel de novo search methods that we present and evaluate here. Standard methods for detecting repetitive sequence are therefore probably missing large numbers of transposable element fragments. In one case (MIR elements), we predict that half of the sequence that is likely present in the genome has gone undetected. Genome-wide, we infer that a large majority of the human genome sequence (>66%–69%) is comprised of repetitive and repeat-derived DNA elements, after controlling for false positives. This estimate stands in stark contrast with previous estimates and suggests that transposable elements have played a much larger role in shaping the history and content of our genome than previously believed.
登录
查看更多内容
影响因子:
4.3
作者:
Li R;Ye J;Li S;Wang J;Han Y;Ye C;Wang J;Yang H;Yu J;Wong GK;Wang J
通讯作者:
Wang J
影响因子:
56.9
作者:
Kirkness, EF;Bafna, V;Venter, JC
通讯作者:
Venter, JC
影响因子:
5.8
作者:
Edgar, RC;Myers, EW
通讯作者:
Myers, EW
影响因子:
7
作者:
Lunter, Gerton;Rocco, Andrea;Hein, Jotun
通讯作者:
Hein, Jotun
影响因子:
3.3
作者:
Castoe TA;Hall KT;Guibotsy Mboulas ML;Gu W;de Koning AP;Fox SE;Poole AW;Vemulapalli V;Daza JM;Mockler T;Smith EN;Feschotte C;Pollock DD
通讯作者:
Pollock DD