Estimating repeat spectra and genome length from low-coverage genome skims with RESPECT.

Estimating repeat spectra and genome length from low-coverage genome skims with RESPECT.
复制标题

DOI:
10.1371/journal.pcbi.1009449
复制
发表时间:
2021-11
影响因子:
4.3
通讯作者:
Bafna V
Bafna V
中科院分区:
生物学2区
文献类型:
--
作者:
Sarmashghi S;Balaban M;Rachtman E;Touri B;Mirarab S;Bafna V

文献摘要

参考文献

被引文献

相似文献

与组装和完成基因组相比,基因组测序的成本下降速度要快得多。使用轻度采样的基因组(基因组撇脂)可能会为基因组生态学带来变革,并且使用 k-mers 的结果表明了这种方法在真核物种的鉴定和系统发育定位方面的优势。在这里,我们重新审视估计基因组参数(例如基因组长度、覆盖度和重复结构)的基本问题,特别关注估计 k 聚体重复谱。我们结合理论和实证分析表明,由于病态系统,估计 k 聚体谱存在根本限制,这对其他基因组参数也有影响。我们使用一种新颖的约束优化方法(样条线性规划)来解决这个问题,其中约束是根据经验学习的。在以 1X 覆盖率模拟 66 个基因组的读数时,我们的方法 REPeat SPECTra Estimation (RESPECT) 的长度估计误差为 2.2%,而之前实现的误差为 27%。在含有污染物的鸟枪法测序读取样本中,RESPECT 长度估计的中值误差为 4%,而其他方法的中值误差为 80%。总之,结果表明低通基因组测序可以对基因组的长度和重复内容进行可靠的估计。 RESPECT 软件将在以下网址公开提供: https://urldefense.proofpoint.com/v2/url?u=https-3A__github.com_shahab-2Dsarmashghi_RESPECT.git&d=DwIGAw&c=-35OiAkTchMrZOngvJPOeA&r=ZozViWvD1E8Por CkfwYKYQMVKFoEcqLFm4Tg49XnPcA&m=f-xS8GMHKckknkc7Xpp8FJYw_ltUwz5frOw1a5pJ8 1EpdTOK8xhbYmrN4ZxniM96&s=717o8hLR1JmHFpRPSWG6xdUQTikyUjicjkipjFsKG4w&e=。与组装和完成基因组相比,基因组测序的成本下降速度要快得多。使用轻度采样的基因组(基因组撇脂)可能会给基因组生态学带来变革。分析基因组撇脂(主要基于小寡聚体的统计数据)仍然具有挑战性,但最近的结果表明这种方法在真核物种的鉴定和系统发育放置方面具有优势。在本文中,我们提出了一种方法“RESPECT”,用于通过低覆盖率基因组略读来估计基因组特性,例如基因组长度和重复性。我们使用组装的基因组来训练 RESPECT,并在低覆盖率的模拟和真实读数上对其进行测试。基准测试结果表明,与其他方法相比,RESPECT 在估计基因组长度方面具有出色的准确性,并且可以提供有关基因组重复结构的关键信息。
The cost of sequencing the genome is dropping at a much faster rate compared to assembling and finishing the genome. The use of lightly sampled genomes (genome-skims) could be transformative for genomic ecology, and results using k-mers have shown the advantage of this approach in identification and phylogenetic placement of eukaryotic species. Here, we revisit the basic question of estimating genomic parameters such as genome length, coverage, and repeat structure, focusing specifically on estimating the k-mer repeat spectrum. We show using a mix of theoretical and empirical analysis that there are fundamental limitations to estimating the k-mer spectra due to ill-conditioned systems, and that has implications for other genomic parameters. We get around this problem using a novel constrained optimization approach (Spline Linear Programming), where the constraints are learned empirically. On reads simulated at 1X coverage from 66 genomes, our method, REPeat SPECTra Estimation (RESPECT), had 2.2% error in length estimation compared to 27% error previously achieved. In shotgun sequenced read samples with contaminants, RESPECT length estimates had median error 4%, in contrast to other methods that had median error 80%. Together, the results suggest that low-pass genomic sequencing can yield reliable estimates of the length and repeat content of the genome. The RESPECT software will be publicly available at https://urldefense.proofpoint.com/v2/url?u=https-3A__github.com_shahab-2Dsarmashghi_RESPECT.git&d=DwIGAw&c=-35OiAkTchMrZOngvJPOeA&r=ZozViWvD1E8PorCkfwYKYQMVKFoEcqLFm4Tg49XnPcA&m=f-xS8GMHKckknkc7Xpp8FJYw_ltUwz5frOw1a5pJ81EpdTOK8xhbYmrN4ZxniM96&s=717o8hLR1JmHFpRPSWG6xdUQTikyUjicjkipjFsKG4w&e=. The cost of sequencing the genome is dropping at a much faster rate compared to assembling and finishing the genome. The use of lightly sampled genomes (genome skims) could be transformative for genomic ecology. Analyzing genome skims, mostly based on statistics of small oligomers, remains challenging, but recent results have shown the advantage of this approach for the identification and phylogenetic placement of eukaryotic species. In this paper, we present a method, RESPECT, to estimate genomic properties such as genome length and repetitiveness from low-coverage genome skims. We trained RESPECT using assembled genomes and tested it on low-coverage simulated and real reads. Benchmarking results reveal that RESPECT has excellent accuracy in estimating the genome length compared to other methods, and can provide critical information regarding the repeat structure of the genome.
DOI: 10.1111/mec.13549
发表时间: 2016-04-01
期刊: MOLECULAR ECOLOGY
影响因子: 4.9
作者:
Coissac, Eric;Hollingsworth, Peter M.;Taberlet, Pierre
通讯作者: Taberlet, Pierre
DOI: 10.1186/s13059-019-1632-4
发表时间: 2019-02-13
期刊: GENOME BIOLOGY
影响因子: 12.3
作者:
Sarmashghi, Shahab;Bohmann, Kristine;Mirarab, Siavash
通讯作者: Mirarab, Siavash
DOI: 10.1186/1471-2105-12-333
发表时间: 2011-08-10
期刊: BMC bioinformatics
影响因子: 3
作者:
Melsted P;Pritchard JK
通讯作者: Pritchard JK
DOI: 10.1093/bioinformatics/btu713
发表时间: 2014-12-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Melsted, Pall;Halldorsson, Bjarni V.
通讯作者: Halldorsson, Bjarni V.
DOI: 10.1098/rspb.2002.2218
发表时间: 2003-02-07
影响因子: 4.7
作者:
Hebert, PDN;Cywinska, A;DeWaard, JR
通讯作者: DeWaard, JR