Using 2k + 2 bubble searches to find SNPs in k-mer graphs

Using 2k + 2 bubble searches to find SNPs in k-mer graphs
复制标题

使用 2k 2 气泡搜索在 k 聚体图中查找 SNP

DOI:
10.1101/004507
复制
发表时间:
2014
期刊:
--
影响因子:
--
通讯作者:
Younsi R
Younsi R
中科院分区:
--
文献类型:
--
作者:
Younsi R

文献摘要

参考文献

相似文献

该预印本现在以出版形式提供:“使用2k+ 2气泡搜索来寻找SNPs ink-mer graphs”,Reda Younsi; Dan MacLean,Bioinformatics 2014; doi:10.1093/bioinformatics/btu 706单核苷酸多态性(SNP)发现是理解遗传变异的重要先决条件。利用目前的测序方法,我们可以全面地对基因组进行采样。通过将序列读数与较长的组装参考比对来发现SNP。De Bruijn图是一种高效的数据结构,可以处理来自现代技术的大量数据。最近的工作表明,这些图的拓扑结构捕获了足够的信息,可以检测和表征遗传变异,为基于比对的方法提供了一种替代方案。这种方法依赖于图的深度优先遍历来识别闭合分叉。这些方法是保守的或产生许多假阳性结果,特别是当遍历图的高度互连(复杂)区域或在非常高的coverage.We设计了一个算法,调用转换的De Bruijn图中的SNPs通过枚举2k+ 2个循环。我们通过与基于比对的方法的SNP列表进行比较来评估预测的SNP的准确性。我们使用来自16个生态型拟南芥的序列数据测试SNP识别的准确性,发现准确性很高。我们发现SNP调用甚至跨越基因组和基因组特征类型。使用图的基于序列的属性来训练决策树使我们能够进一步提高SNP调用的准确性。这些结果共同表明,我们的算法能够在复杂的子图中准确地找到SNP,并且可能全面地从全基因组图中找到SNP。我们算法的C++实现的源代码可以在GNU公共许可证v3下获得:https://github.com/redayounsi/2kplus2
This preprint is now available in published form as: ‘Using 2k+ 2 bubble searches to find SNPs ink-mer graphs’, Reda Younsi; Dan MacLean, Bioinformatics 2014; doi: 10.1093/bioinformatics/btu706Single Nucleotide Polymorphism (SNP) discovery is an important preliminary for understanding genetic variation. With current sequencing methods we can sample genomes comprehensively. SNPs are found by aligning sequence reads against longer assembled references. De Bruijn graphs are efficient data structures that can deal with the vast amount of data from modern technologies. Recent work has shown that the topology of these graphs captures enough information to allow the detection and characterisation of genetic variants, offering an alternative to alignment-based methods. Such methods rely on depth-first walks of the graph to identify closing bifurcations. These methods are conservative or generate many false-positive results, particularly when traversing highly inter-connected (complex) regions of the graph or in regions of very high coverage.We devised an algorithm that calls SNPs in converted De Bruijn graphs by enumerating 2k+ 2 cycles. We evaluated the accuracy of predicted SNPs by comparison with SNP lists from alignment based methods. We tested accuracy of the SNP calling using sequence data from sixteen ecotypes ofArabidopsis thalianaand found that accuracy was high. We found that SNP calling was even across the genome and genomic feature types. Using sequence based attributes of the graph to train a decision tree allowed us to increase accuracy of SNP calls further.Together these results indicate that our algorithm is capable of finding SNPs accurately in complex sub-graphs and potentially comprehensively from whole genome graphs.The source code for a C++ implementation of our algorithm is available under the GNU Public Licence v3 at:https://github.com/redayounsi/2kplus2
DOI: 10.1016/s0022-2836(05)80360-2
发表时间: 1990-10-05
影响因子: 5.6
作者:
ALTSCHUL, SF;GISH, W;LIPMAN, DJ
通讯作者: LIPMAN, DJ