From telomere to telomere: The transcriptional and epigenetic state of human repeat elements.

From telomere to telomere: The transcriptional and epigenetic state of human repeat elements.
复制标题

DOI:
10.1126/science.abk3112
复制
发表时间:
2022-04
期刊:
影响因子:
56.9
通讯作者:
O'Neill, Rachel J.
O'Neill, Rachel J.
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Hoyt, Savannah J.;Storer, Jessica M.;Hartley, Gabrielle A.;Grady, Patrick G. S.;Gershman, Ariel;de Lima, Leonardo G.;Limouse, Charles;Halabian, Reza;Wojenski, Luke;Rodriguez, Matias;Altemose, Nicolas;Rhie, Arang;Core, Leighton J.;Gerton, Jennifer L.;Makalowski, Wojciech;Olson, Daniel;Rosen, Jeb;Smit, Arian F. A.;Straight, Aaron F.;Vollger, Mitchell R.;Wheeler, Travis J.;Schatz, Michael C.;Eichler, Evan E.;Phillippy, Adam M.;Timp, Winston;Miga, Karen H.;O'Neill, Rachel J.

文献摘要

参考文献

被引文献

相似文献

移动元件和重复基因组区域是谱系特异性基因组创新和独特的个体基因组指纹的来源。对此类重复元件(包括在基因组更复杂区域中发现的重复元件)的综合分析需要完整的线性基因组组装。我们提出了 T2T-CHM13 人类参考基因组的从头重复发现和注释。我们鉴定了以前未知的卫星阵列,扩展了重复和移动元件的变体和家族目录,表征了复杂复合重复的类别,并定位了逆转录元件转导事件。我们检测了新生转录并描绘了 CpG 甲基化谱,以定义人类转录活性逆转录元件的结构,包括着丝粒中的逆转录元件。这些数据扩展了我们对塑造人类基因组的重复区域的多样性、分布和进化的了解。 CHM13 的端粒到端粒组装支持重复注释和发现。人类参考 T2T-CHM13 填补了 GRCh38 中的空白并纠正了塌陷区域(三角形)。结合基于长读的甲基化调用、PRO-seq 和多级计算方法,我们提供了人类重复序列的概要,定义了逆转录元件表达和甲基化图谱,并描绘了全基因组新生转录的位点特异性位点,包括以前无法访问的着丝粒。 SINE,短散布单元; SVA、SINE——可变数量串联重复序列——Alu; LINE,长散布元素; LTR,长末端重复; TSS,转录起始位点; pA,聚腺苷酸化信号。转座元件 (TE)、重复扩增和重复介导的结构重排在染色体结构和物种进化中发挥着关键作用,有助于人类遗传变异,并通过拷贝数变异、结构变异、插入、缺失以及基因转录和剪接的改变极大地影响人类健康。尽管重复区域在基因组稳定性中发挥着重要作用,但由于其开发过程中的技术限制,重复区域已被归为人类基因组参考 GRCh38 中的间隙和折叠区域。这些区域(特别是着丝粒)缺乏线性序列,导致无法在局部和区域染色体环境中充分探索人类基因组的重复内容。长读长测序支持伪单倍体人类细胞系 CHM13 的完整端粒到端粒 (T2T) 组装。该资源提供了对所有人类重复序列的基因组规模评估,包括 TE 和以前未知的重复序列和卫星,无论是在间隙和折叠区域内部还是外部。此外,完整的基因组使我们有机会探索这些元件的表观遗传和转录谱,这对于我们理解染色体结构、功能和进化至关重要。比较分析揭示了重复分歧、进化、扩展或收缩的模式以及位点级分辨率。我们使用先前已知的人类重复和从头重复建模,然后进行手动管理,实现了全面的重复注释工作流程,包括评估基因注释的重叠、片段重复、串联重复和注释重复。使用这种方法,我们开发了人类重复序列的更新目录,并完善了之前的重复注释。我们在 T2T-CHM13 中发现了 43 个以前未知的重复和重复变体,并表征了 19 个复杂的复合重复结构,这些结构通常携带基因。使用精确核连续测序 (PRO-seq) 和从 Oxford Nanopore Technologies 长读长测序数据生成的 CpG 甲基化位点,我们评估了全基因组范围内逆转录元件的 RNA 聚合酶参与,揭示了新生转录、序列分歧、CpG 密度和甲基化之间的相关性。这些分析扩展到评估所有重复序列的 RNA 聚合酶占用率,包括位于所有人类染色体先前无法到达的着丝粒区域的高密度卫星重复序列。此外,在早期发育阶段和完整的细胞周期时间序列中使用图谱依赖和图谱独立的方法,我们发现跨卫星的RNA聚合酶参与度较低;相比之下,TE 转录丰富,可作为 CpG 甲基化和着丝粒亚结构变化的边界。这些数据共同揭示了转录活性逆转录元件亚类与 DNA 甲基化之间的动态关系,以及新重复家族和复合元件的衍生和进化的潜在机制。着眼于 HG002 X 染色体的新兴 T2T 水平组装,我们发现人类群体中可能存在高水平的重复变异,包括影响基因拷贝数的复合元件拷贝数。此外,我们强调了重复对基因组结构多样性的影响,揭示了人类和灵长类动物之间具有极端拷贝数差异的重复扩展,同时还提供了逆转录元件转导事件的高可信度注释。本文描述的全面重复注释和更新的重复模型可作为扩展人类基因组序列纲要的资源,并揭示特定重复对人类基因组的影响。在开发该资源时,我们提供了一个方法框架,用于评估人类基因组内部和之间的重复变异。在基因组规模和局部(例如着丝粒内)对重复的转录景观进行详尽的评估,为功能研究奠定了基础,以阐明转录在基因组稳定性和染色体分离所必需的机制中所扮演的角色。最后,我们的工作表明需要加大力度实现非人类灵长类动物和其他物种的 T2T 水平组装,以充分了解定义灵长类动物谱系(包括人类)的重复衍生基因组创新的复杂性和影响。
Mobile elements and repetitive genomic regions are sources of lineage-specific genomic innovation and uniquely fingerprint individual genomes. Comprehensive analyses of such repeat elements, including those found in more complex regions of the genome, require a complete, linear genome assembly. We present a de novo repeat discovery and annotation of the T2T-CHM13 human reference genome. We identified previously unknown satellite arrays, expanded the catalog of variants and families for repeats and mobile elements, characterized classes of complex composite repeats, and located retroelement transduction events. We detected nascent transcription and delineated CpG methylation profiles to define the structure of transcriptionally active retroelements in humans, including those in centromeres. These data expand our insight into the diversity, distribution, and evolution of repetitive regions that have shaped the human genome. Telomere-to-telomere assembly of CHM13 supports repeat annotations and discoveries. The human reference T2T-CHM13 filled gaps and corrected collapsed regions (triangles) in GRCh38. Combining long read–based methylation calls, PRO-seq, and multilevel computational methods, we provide a compendium of human repeats, define retroelement expression and methylation profiles, and delineate locus-specific sites of nascent transcription genome-wide, including previously inaccessible centromeres. SINE, short interspersed element; SVA, SINE–variable number tandem repeat–Alu; LINE, long interspersed element; LTR, long terminal repeat; TSS, transcription start site; pA, polyadenylation signal. Transposable elements (TEs), repeat expansions, and repeat-mediated structural rearrangements play key roles in chromosome structure and species evolution, contribute to human genetic variation, and substantially influence human health through copy number variants, structural variants, insertions, deletions, and alterations to gene transcription and splicing. Despite their formative role in genome stability, repetitive regions have been relegated to gaps and collapsed regions in human genome reference GRCh38 owing to the technological limitations during its development. The lack of linear sequence in these regions, particularly in centromeres, resulted in the inability to fully explore the repeat content of the human genome in the context of both local and regional chromosomal environments. Long-read sequencing supported the complete, telomere-to-telomere (T2T) assembly of the pseudo-haploid human cell line CHM13. This resource affords a genome-scale assessment of all human repetitive sequences, including TEs and previously unknown repeats and satellites, both within and outside of gaps and collapsed regions. Additionally, a complete genome enables the opportunity to explore the epigenetic and transcriptional profiles of these elements that are fundamental to our understanding of chromosome structure, function, and evolution. Comparative analyses reveal modes of repeat divergence, evolution, and expansion or contraction with locus-level resolution. We implemented a comprehensive repeat annotation workflow using previously known human repeats and de novo repeat modeling followed by manual curation, including assessing overlaps with gene annotations, segmental duplications, tandem repeats, and annotated repeats. Using this method, we developed an updated catalog of human repetitive sequences and refined previous repeat annotations. We discovered 43 previously unknown repeats and repeat variants and characterized 19 complex, composite repetitive structures, which often carry genes, across T2T-CHM13. Using precision nuclear run-on sequencing (PRO-seq) and CpG methylated sites generated from Oxford Nanopore Technologies long-read sequencing data, we assessed RNA polymerase engagement across retroelements genome-wide, revealing correlations between nascent transcription, sequence divergence, CpG density, and methylation. These analyses were extended to evaluate RNA polymerase occupancy for all repeats, including high-density satellite repeats that reside in previously inaccessible centromeric regions of all human chromosomes. Moreover, using both mapping-dependent and mapping-independent approaches across early developmental stages and a complete cell cycle time series, we found that engaged RNA polymerase across satellites is low; in contrast, TE transcription is abundant and serves as a boundary for changes in CpG methylation and centromere substructure. Together, these data reveal the dynamic relationship between transcriptionally active retroelement subclasses and DNA methylation, as well as potential mechanisms for the derivation and evolution of new repeat families and composite elements. Focusing on the emerging T2T-level assembly of the HG002 X chromosome, we reveal that a high level of repeat variation likely exists across the human population, including composite element copy numbers that affect gene copy number. Additionally, we highlight the impact of repeats on the structural diversity of the genome, revealing repeat expansions with extreme copy number differences between humans and primates while also providing high-confidence annotations of retroelement transduction events. The comprehensive repeat annotations and updated repeat models described herein serve as a resource for expanding the compendium of human genome sequences and reveal the impact of specific repeats on the human genome. In developing this resource, we provide a methodological framework for assessing repeat variation within and between human genomes. The exhaustive assessment of the transcriptional landscape of repeats, at both the genome scale and locally, such as within centromeres, sets the stage for functional studies to disentangle the role transcription plays in the mechanisms essential for genome stability and chromosome segregation. Finally, our work demonstrates the need to increase efforts toward achieving T2T-level assemblies for nonhuman primates and other species to fully understand the complexity and impact of repeat-derived genomic innovations that define primate lineages, including humans.
DOI: 10.1038/nbt.1533
发表时间: 2009-04
影响因子: 46.9
作者:
Ball, Madeleine P.;Li, Jin Billy;Gao, Yuan;Lee, Je-Hyuk;LeProust, Emily M.;Park, In-Hyun;Xie, Bin;Daley, George Q.;Church, George M.
通讯作者: Church, George M.
DOI: 10.1371/journal.pgen.1000354
发表时间: 2009-01
期刊: PLOS GENETICS
影响因子: 4.5
作者:
Chueh, Anderly C.;Northrop, Emma L.;Brettingham-Moore, Kate H.;Choo, K. H. Andy;Wong, Lee H.
通讯作者: Wong, Lee H.
DOI: 10.1016/j.devcel.2015.05.012
发表时间: 2015-07-06
期刊: DEVELOPMENTAL CELL
影响因子: 11.8
作者:
Chen, Chin-Chi;Bowers, Sarion;Lipinszki, Zoltan;Palladino, Jason;Trusiak, Sarah;Bettini, Emily;Rosin, Leah;Przewloka, Marcin R.;Glover, David M.;O'Neill, Rachel J.;Mellone, Barbara G.
通讯作者: Mellone, Barbara G.
DOI: 10.1146/annurev-genom-082509-141802
发表时间: 2011
影响因子: 8.7
作者:
Beck CR;Garcia-Perez JL;Badge RM;Moran JV
通讯作者: Moran JV
DOI: 10.1097/pgp.0000000000000697
发表时间: 2021-07-01
期刊: International journal of gynecological pathology : official journal of the International Society of Gynecological Pathologists
影响因子: --
作者:
Chen X;Ma Y;Wang L;Zhang X;Yu Y;Lü W;Xie X;Cheng X
通讯作者: Cheng X