Semi-automated assembly of high-quality diploid human reference genomes.

Semi-automated assembly of high-quality diploid human reference genomes.
复制标题

DOI:
10.1038/s41586-022-05325-5
复制
发表时间:
2022-11
期刊:
影响因子:
64.8
通讯作者:
Miga, Karen H
Miga, Karen H
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Jarvis, Erich D;Formenti, Giulio;Rhie, Arang;Guarracino, Andrea;Yang, Chentao;Wood, Jonathan;Tracey, Alan;Thibaud-Nissen, Francoise;Vollger, Mitchell R;Porubsky, David;Cheng, Haoyu;Asri, Mobin;Logsdon, Glennis A;Carnevali, Paolo;Chaisson, Mark J P;Chin, Chen-Shan;Cody, Sarah;Collins, Joanna;Ebert, Peter;Escalona, Merly;Fedrigo, Olivier;Fulton, Robert S;Fulton, Lucinda L;Garg, Shilpa;Gerton, Jennifer L;Ghurye, Jay;Granat, Anastasiya;Green, Richard E;Harvey, William;Hasenfeld, Patrick;Hastie, Alex;Haukness, Marina;Jaeger, Erich B;Jain, Miten;Kirsche, Melanie;Kolmogorov, Mikhail;Korbel, Jan O;Koren, Sergey;Korlach, Jonas;Lee, Joyce;Li, Daofeng;Lindsay, Tina;Lucas, Julian;Luo, Feng;Marschall, Tobias;Mitchell, Matthew W;McDaniel, Jennifer;Nie, Fan;Olsen, Hugh E;Olson, Nathan D;Pesout, Trevor;Potapova, Tamara;Puiu, Daniela;Regier, Allison;Ruan, Jue;Salzberg, Steven L;Sanders, Ashley D;Schatz, Michael C;Schmitt, Anthony;Schneider, Valerie A;Selvaraj, Siddarth;Shafin, Kishwar;Shumate, Alaina;Stitziel, Nathan O;Stober, Catherine;Torrance, James;Wagner, Justin;Wang, Jianxin;Wenger, Aaron;Xiao, Chuanle;Zimin, Aleksey V;Zhang, Guojie;Wang, Ting;Li, Heng;Garrison, Erik;Haussler, David;Hall, Ira;Zook, Justin M;Eichler, Evan E;Phillippy, Adam M;Paten, Benedict;Howe, Kerstin;Miga, Karen H

文献摘要

被引文献

相似文献

目前的人类参考基因组GRCh38代表了20多年的努力,以产生高质量的组装,这使社会受益。然而,它仍然有许多差距和错误,并且不代表生物基因组,因为它是多个个体的混合物。最近,一个高质量的端粒到端粒的参考,CHM13,产生了最新的长读技术,但它是来自一个葡萄胎细胞系与近纯合的基因组。为了解决这些局限性,人类泛基因组参考联盟成立,目标是为代表人类遗传多样性的泛基因组参考创建高质量,具有成本效益的二倍体基因组组装。在这里,在我们的第一份科学报告中,我们确定了当前基因组测序和组装方法的哪种组合以最少的人工管理产生最完整和准确的二倍体基因组组装。在组装期间使用高度准确的长读段和具有基于图的单倍型定相的父子数据的方法优于那些没有的方法。通过开发最佳性能方法的组合,我们生成了第一个高质量的二倍体参考组装体,平均每条染色体仅包含大约四个缺口,大多数染色体在CHM 13长度的±1%范围内。近48%的蛋白质编码基因在单倍型之间存在非同义氨基酸变化,其中着丝粒区域的多样性最高。我们的研究结果为大规模组装接近完整的二倍体人类基因组奠定了基础,以作为泛基因组参考,以捕获从单核苷酸到结构重排的全球遗传变异。确定当前基因组测序和组装方法的哪种组合导致高质量的完整二倍体基因组组装。
The current human reference genome, GRCh38, represents over 20 years of effort to generate a high-quality assembly, which has benefitted society. However, it still has many gaps and errors, and does not represent a biological genome as it is a blend of multiple individuals. Recently, a high-quality telomere-to-telomere reference, CHM13, was generated with the latest long-read technologies, but it was derived from a hydatidiform mole cell line with a nearly homozygous genome. To address these limitations, the Human Pangenome Reference Consortium formed with the goal of creating high-quality, cost-effective, diploid genome assemblies for a pangenome reference that represents human genetic diversity. Here, in our first scientific report, we determined which combination of current genome sequencing and assembly approaches yield the most complete and accurate diploid genome assembly with minimal manual curation. Approaches that used highly accurate long reads and parent–child data with graph-based haplotype phasing during assembly outperformed those that did not. Developing a combination of the top-performing methods, we generated our first high-quality diploid reference assembly, containing only approximately four gaps per chromosome on average, with most chromosomes within ±1% of the length of CHM13. Nearly 48% of protein-coding genes have non-synonymous amino acid changes between haplotypes, and centromeric regions showed the highest diversity. Our findings serve as a foundation for assembling near-complete diploid human genomes at scale for a pangenome reference to capture global genetic variation from single nucleotides to structural rearrangements. Which combination of current genome sequencing and assembly approaches results in high-quality, complete diploid genome assemblies is determined.