Measuring the invisible - The sequences causal of genome size differences in eyebrights ( Euphrasia ) revealed by k-mers

Measuring the invisible - The sequences causal of genome size differences in eyebrights ( Euphrasia ) revealed by k-mers
复制标题

测量不可见的东西 - k-mers 揭示的导致眼球(Euphasia)基因组大小差异的序列

DOI:
10.1101/2021.11.09.467866
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
Becher H
Becher H
中科院分区:
--
文献类型:
--
作者:
Becher H

文献摘要

相似文献

植物类群内基因组大小的变异是由于存在/不存在变异,这可能会影响各种频率类别的低拷贝序列或基因组重复序列。然而,识别基因组大小变异的序列是具有挑战性的,因为基因组组装通常包含重复序列的折叠表示,并且因为基因组撇取研究设计错过低拷贝数序列。在这里,我们采取了一种新的方法的基础上k-mer,短的子序列的长度相等k,产生的全基因组测序数据的二倍体eyebrights(Euphrasia),一组植物,具有相当大的基因组大小的变化倍性水平。我们比较的k-mer库存内和密切相关的物种之间,并量化不同的拷贝数类基因组大小差异的贡献。我们进一步将高拷贝数k-mer与从RepeatExplorer 2管道中检索到的特定重复类型进行匹配。我们发现基因组大小差异高达230 Mbp,相当于超过20%的基因组大小变异。这些差异的最大贡献来自rDNA序列,一个145-nt的基因组卫星和重复与安吉拉转座因子。我们还发现低拷贝数类别(拷贝数≤ 10×)的大小差异高达27 Mbp,这可能表明我们样本之间的基因空间差异。我们证明,它是可能的,以查明序列引起的基因组大小变化的物种内,而不使用参考基因组。这样的序列可以作为未来细胞遗传学研究的目标。我们还表明,基因组大小变异的研究应该超越重复,如果他们的目标是确定全方位的基因组变异。为了允许将来与其他分类组的工作,我们分享了我们的k-mer分析管道,它很容易运行,主要依赖于标准的GNU命令行工具。
Genome size variation within plant taxa is due to presence/absence variation, which may affect low-copy sequences or genomic repeats of various frequency classes. However, identifying the sequences underpinning genome size variation is challenging because genome assemblies commonly contain collapsed representations of repetitive sequences and because genome skimming studies by design miss low-copy number sequences. Here, we take a novel approach based on k-mers, short sub-sequences of equal lengthk, generated from whole-genome sequencing data of diploid eyebrights (Euphrasia), a group of plants that have considerable genome size variation within a ploidy level. We compare k-mer inventories within and between closely related species, and quantify the contribution of different copy number classes to genome size differences. We further match high-copy number k-mers to specific repeat types as retrieved from the RepeatExplorer2 pipeline. We find genome size differences of up to 230Mbp, equivalent to more than 20% genome size variation. The largest contributions to these differences come from rDNA sequences, a 145-nt genomic satellite and a repeat associated with an Angela transposable element. We also find size differences in the low-copy number class (copy number ≤ 10×) of up to 27 Mbp, possibly indicating differences in gene space between our samples. We demonstrate that it is possible to pinpoint the sequences causing genome size variation within species without the use of a reference genome. Such sequences can serve as targets for future cytogenetic studies. We also show that studies of genome size variation should go beyond repeats if they aim to characterise the full range of genomic variants. To allow future work with other taxonomic groups, we share our k-mer analysis pipeline, which is straightforward to run, relying largely on standard GNU command line tools.