An accurate algorithm for the detection of DNA fragments from dilution pool sequencing experiments.

An accurate algorithm for the detection of DNA fragments from dilution pool sequencing experiments.
复制标题

用于检测稀释池测序实验中 DNA 片段的准确算法。

DOI:
10.1093/bioinformatics/btx436
复制
发表时间:
2018
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Bansal,Vikas
Bansal,Vikas
中科院分区:
--
文献类型:
--
作者:
Bansal,Vikas

文献摘要

相似文献

目前高通量测序技术的短读取长度限制了恢复远距离单倍型信息的能力。从高分子量DNA片段制备DNA测序文库的稀释池方法使从短序列读取中回收长DNA片段成为可能。这些方法需要计算方法,使用比对的序列读取来识别DNA片段,并将片段组装成长的单倍型。结果我们将稀释池测序实验中DNA片段的检测问题描述为一个基因组分割问题,并提出了一种算法,该算法利用动态规划来优化序列读取的生成模型中的似然函数。该算法使用迭代方法自动推断平均背景读取深度和每个池中的碎片数量。使用模拟数据,我们证明了我们的方法FragmentCut比基于HMM的片段检测方法具有25%-30%的灵敏度,并且还可以检测出重叠的片段。在一个全基因组人类融合体库数据集上,与现有的两种方法相比,使用FragmentCut鉴定的片段组装的单倍型具有更长的N50长度,切换错误率降低16.2%,错配错误率降低35.8%。我们使用另外两个稀释池数据集进一步证明了我们的方法的更高的准确性。可用性和实现从https://bansal-lab.github.io/software/FragmentCutSupplementary信息中获得片段切割补充数据可以在BioInformation在线上获得。
MotivationThe short read lengths of current high-throughput sequencing technologies limit the ability to recover long-range haplotype information. Dilution pool methods for preparing DNA sequencing libraries from high molecular weight DNA fragments enable the recovery of long DNA fragments from short sequence reads. These approaches require computational methods for identifying the DNA fragments using aligned sequence reads and assembling the fragments into long haplotypes. Although a number of computational methods have been developed for haplotype assembly, the problem of identifying DNA fragments from dilution pool sequence data has not received much attention.ResultsWe formulate the problem of detecting DNA fragments from dilution pool sequencing experiments as a genome segmentation problem and develop an algorithm that uses dynamic programming to optimize a likelihood function derived from a generative model for the sequence reads. This algorithm uses an iterative approach to automatically infer the mean background read depth and the number of fragments in each pool. Using simulated data, we demonstrate that our method, FragmentCut, has 25–30% greater sensitivity compared with an HMM based method for fragment detection and can also detect overlapping fragments. On a whole-genome human fosmid pool dataset, the haplotypes assembled using the fragments identified by FragmentCut had greater N50 length, 16.2% lower switch error rate and 35.8% lower mismatch error rate compared with two existing methods. We further demonstrate the greater accuracy of our method using two additional dilution pool datasets.Availability and implementationFragmentCut is available from https://bansal-lab.github.io/software/FragmentCutSupplementary informationSupplementary data are available atBioinformaticsonline.