GapReduce: A Gap Filling Algorithm Based on Partitioned Read Sets

GapReduce: A Gap Filling Algorithm Based on Partitioned Read Sets
复制标题

DOI:
10.1109/tcbb.2018.2789909
复制
发表时间:
2020-05
期刊:
IEEE/ACM Transactions on Computational Biology and Bioinformatics
影响因子:
--
通讯作者:
Junwei Luo;Jianxin Wang;Juan Shang;Huimin Luo;Min Li;Fangxiang Wu;Yi Pan
Junwei Luo;Jianxin Wang;Juan Shang;Huimin Luo;Min Li;Fangxiang Wu;Yi Pan
中科院分区:
其他
文献类型:
--
作者:
Junwei Luo;Jianxin Wang;Juan Shang;Huimin Luo;Min Li;Fangxiang Wu;Yi Pan

文献摘要

相似文献

随着测序和组装技术的进步,越来越多基因组的草图序列可供使用。然而,这些草图序列中通常存在缺口,这会影响生物学研究的各种下游分析。缺口填补方法可以缩短缺口长度并提高这些基因组草图序列的完整性。尽管已经开发了一些缺口填补工具,但它们的有效性和准确性仍需提高。在这项研究中,我们开发了一种名为GapReduce的新工具,它可以使用成对读段来填补缺口。对于一个缺口,GapReduce选择其配对读段在左侧或右侧侧翼区域比对的读段,并将这些读段划分为两组。然后,GapReduce采用不同的\(k\)值和\(k\)-mer频率阈值来迭代构建德布鲁因图,这些图用于找到填补缺口的正确路径。为了克服在路径选择过程中由重复区域和测序错误导致的分支问题,GapReduce设计了一种新方法,该方法基于划分的读段集同时考虑\(k\)-mer频率和成对读段的分布。我们将GapReduce的性能与当前流行的缺口填补工具进行了比较。实验结果表明,GapReduce能够产生令人满意的缺口填补结果,特别是对于长插入片段大小的数据集。GapReduce可在https://github.com/bioinfomaticsCSU/GapReduce公开下载。
With the advances in technologies of sequencing and assembly, draft sequences of more and more genomes are available. However, there commonly exist gaps in these draft sequences which influence various downstream analysis of biological studies. Gap filling methods can shorten the length of gaps and improve the completion of these draft sequences of genomes. Although some gap filling tools have been developed, their effectiveness and accuracy need to be improved. In this study, we develop a novel tool, called GapReduce, which can fill the gaps using the paired reads. For a gap, GapReduce selects the reads whose mate reads are aligned on the left or the right flanking region, and partitions the reads to two sets. Then GapReduce adopts different <inline-formula><tex-math notation="LaTeX">$k$</tex-math><alternatives><mml:math><mml:mi>k</mml:mi></mml:math><inline-graphic xlink:href="wang-ieq1-2789909.gif"/></alternatives></inline-formula> values and <inline-formula><tex-math notation="LaTeX">$k$</tex-math><alternatives><mml:math><mml:mi>k</mml:mi></mml:math><inline-graphic xlink:href="wang-ieq2-2789909.gif"/></alternatives></inline-formula>-<inline-formula><tex-math notation="LaTeX">$mer$</tex-math><alternatives><mml:math><mml:mrow><mml:mi>m</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:math><inline-graphic xlink:href="wang-ieq3-2789909.gif"/></alternatives></inline-formula> frequency thresholds to iteratively construct De Bruijn graphs, which are used for finding the correct path to fill the gap. For overcoming the branching problems caused by repetitive regions and sequencing errors in the procedure of path selection, GapReduce designs a novel approach that simultaneously considers <inline-formula><tex-math notation="LaTeX">$k$</tex-math><alternatives><mml:math><mml:mi>k</mml:mi></mml:math><inline-graphic xlink:href="wang-ieq4-2789909.gif"/></alternatives></inline-formula>-<inline-formula><tex-math notation="LaTeX">$mer$</tex-math><alternatives><mml:math><mml:mrow><mml:mi>m</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:math><inline-graphic xlink:href="wang-ieq5-2789909.gif"/></alternatives></inline-formula> frequency and distribution of paired reads based on the partitioned read sets. We compare the performance of GapReduce with current popular gap filling tools. The experimental results demonstrate that GapReduce can produce satisfactory gap filling results, especially for long insert size datasets. GapReduce is publicly available for downloading at <uri>https://github.com/bioinfomaticsCSU/GapReduce</uri>.