Probably Correct: Rescuing Repeats with Short and Long Reads.

Probably Correct: Rescuing Repeats with Short and Long Reads.
复制标题

DOI:
10.3390/genes12010048
复制
发表时间:
2020-12-31
期刊:
影响因子:
3.5
通讯作者:
Cechova M
Cechova M
中科院分区:
生物学3区
文献类型:
--
作者:
Cechova M

文献摘要

参考文献

被引文献

相似文献

自从在人类基因组计划之后引入高通量测序以来,将短的读数组装成足够高质量的参照物构成了一个重大问题,因为人类基因组的很大一部分--估计为50%-69%--是重复的。因此,相当大比例的测序读数是多图谱的,即在基因组中没有唯一的位置。读取是否为多重映射的两个关键参数是读取长度和基因组复杂性。长阅读现在能够跨越困难的异染色质区域,包括完整的着丝粒,并表征从“端粒到端粒”的染色体。此外,相同的读取或重复阵列可以根据它们的表观遗传标记来区分,例如甲基化模式,这有助于组装过程。这是尽管长读数仍然含有一定比例的测序错误,这在准确性和速度上都让比对者和组装者迷失了方向。在这里,我回顾了针对重复解决和多重映射读取问题的建议和实施的解决方案,以及参考选择、重复掩蔽和性染色体的适当表示所产生的下游后果。我还考虑了即将到来的与长阅读有关的挑战和解决方案,我们预计将从单个个体内的重复定位问题转变为Pangenome内的重复定位问题。
Ever since the introduction of high-throughput sequencing following the human genome project, assembling short reads into a reference of sufficient quality posed a significant problem as a large portion of the human genome—estimated 50–69%—is repetitive. As a result, a sizable proportion of sequencing reads is multi-mapping, i.e., without a unique placement in the genome. The two key parameters for whether or not a read is multi-mapping are the read length and genome complexity. Long reads are now able to span difficult, heterochromatic regions, including full centromeres, and characterize chromosomes from “telomere to telomere”. Moreover, identical reads or repeat arrays can be differentiated based on their epigenetic marks, such as methylation patterns, aiding in the assembly process. This is despite the fact that long reads still contain a modest percentage of sequencing errors, disorienting the aligners and assemblers both in accuracy and speed. Here, I review the proposed and implemented solutions to the repeat resolution and the multi-mapping read problem, as well as the downstream consequences of reference choice, repeat masking, and proper representation of sex chromosomes. I also consider the forthcoming challenges and solutions with regards to long reads, where we expect the shift from the problem of repeat localization within a single individual to the problem of repeat positioning within pangenomes.
DOI: 10.1186/s13742-015-0052-y
发表时间: 2015
期刊: GigaScience
影响因子: 9.2
作者:
Howe K;Wood JM
通讯作者: Wood JM
DOI: 10.1038/s41588-020-0671-9
发表时间: 2020-07-27
期刊: NATURE GENETICS
影响因子: 30.8
作者:
Haberer, Georg;Kamal, Nadia;Mayer, Klaus F. X.
通讯作者: Mayer, Klaus F. X.
DOI: 10.1371/journal.pgen.1002384
发表时间: 2011-12
期刊: PLoS genetics
影响因子: 4.5
作者:
de Koning AP;Gu W;Castoe TA;Batzer MA;Pollock DD
通讯作者: Pollock DD
DOI: 10.1073/pnas.2001749117
发表时间: 2020-10-20
影响因子: 11.1
作者:
Cechova M;Vegesna R;Tomaszkiewicz M;Harris RS;Chen D;Rangavittal S;Medvedev P;Makova KD
通讯作者: Makova KD
DOI: 10.1093/molbev/msz156
发表时间: 2019-11-01
影响因子: 10.7
作者:
Cechova, Monika;Harris, Robert S.;Makova, Kateryna D.
通讯作者: Makova, Kateryna D.