Evaluating the use of variation graphs for the characterisation of yeast rDNA arrays

Evaluating the use of variation graphs for the characterisation of yeast rDNA arrays
复制标题

评估变异图在酵母 rDNA 阵列表征中的使用

DOI:
10.1099/acmi.byg2019.po0016
复制
发表时间:
2019
影响因子:
--
通讯作者:
Ursani Z
Ursani Z
中科院分区:
--
文献类型:
--
作者:
Ursani Z

文献摘要

相似文献

在过去的十年中,我们已经取得了相当大的洞察力,以确定的rDNA阵列内的序列变异的Saccharomycesparadoxus和其最近的野生亲戚。然而,相当大的挑战仍然存在于这个复杂的基因组区域的计算表征。本研究旨在评估变异图在这方面的应用,正式比较它们与传统线性方法的有效性。具体而言,我们的目的是在10个不同的、具有高质量基因组数据集的单倍体酿酒酵母菌株的rDNA阵列中识别部分和固定变异(即pSNP、SNP、pINDEL和INDEL)。我们使用两种截然不同的方法构建了两个计算管道。第一个流水线使用BWA读取映射器和BCFtools变体调用器来识别针对线性S288 c参考的变体,第二个流水线使用vg工具来调用针对图形参考的变体结果表明,基于图形的流水线能够比线性流水线识别更多的变体,特别是部分变异,同时也遗漏了BWA/BCF工具识别的一些关键变异。在vg管道鉴定变体的基因座处的读段覆盖率中发现了两个管道之间的主要差异。在接下来的几个月里,我们的目标是调查这些差异的原因,并开发一种新的基于图形的计算管道,可以准确地识别这个关键基因组区域内的全范围序列和拷贝数变异。
Over the past decade, we have gained considerable insight into the identification of sequence variation within the rDNA array ofSaccharomyces cerevisiaeand its closest wild relative,Saccharomyces paradoxus. Yet considerable challenges remain in the computational characterisation of this complex genomic region. This study aimed to evaluate the use of variation graphs for this purpose, formally comparing their effectiveness with traditional linear approaches.Specifically, we aimed to identify both partial and fixed variants (i.e. pSNPs, SNPs, pINDELs and INDELs) in the rDNA arrays of 10 diverse, haploidSaccharomyces cerevisiaestrains with high quality genomic datasets. We constructed two computational pipelines using two highly different approaches. The first pipeline used the BWA read mapper and the BCFtools variant caller to identify variants against the linear S288c reference, with the second pipeline using the vg tool to call variants against a graphical reference (either based on a graphical representation of the S288c genome or aSaccharomyces cerevisiaepan-genome).The results showed that the graph-based pipeline was able to identify more variants than the linear pipeline, and in particular partial variants, while also missing some key variants identified by BWA/BCFtools. A major discrepancy between the two pipelines was found in the read coverage at loci where the vg pipeline identified variants. In the coming months, we aim to investigate the cause of these differences and to develop a new graph-based computational pipeline that can accurately identify the full range of sequence and copy number variation within this key genomic region.