Large Scale Loss of Data in Low-Diversity Illumina Sequencing Libraries Can Be Recovered by Deferred Cluster Calling

Large Scale Loss of Data in Low-Diversity Illumina Sequencing Libraries Can Be Recovered by Deferred Cluster Calling
复制标题

DOI:
10.1371/journal.pone.0016607
复制
发表时间:
2011-01-28
期刊:
影响因子:
3.7
通讯作者:
Osborne, Cameron S.
Osborne, Cameron S.
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Krueger, Felix;Andrews, Simon R.;Osborne, Cameron S.

文献摘要

被引文献

相似文献

大规模并行DNA测序能够同时对数千万个DNA片段进行测序。然而,用于确定单个簇的坐标的初始循环中的序列偏差导致Illumina基因组分析仪上的簇鉴定的保真度损失。这可以导致可以分析的聚类数量的显著减少。这种低样品多样性是通过限制性酶消化产生的测序文库(例如e4 C-seq或简化代表文库)的固有问题。类似地,这个问题也可能通过条形码化的多重文库的组合测序而出现。我们描述了一个程序,以推迟集群坐标的映射,直到低多样性序列已经通过。这个简单的程序可以恢复大量的下一代测序数据,否则将丢失。
Massively parallel DNA sequencing is capable of sequencing tens of millions of DNA fragments at the same time. However, sequence bias in the initial cycles, which are used to determine the coordinates of individual clusters, causes a loss of fidelity in cluster identification on Illumina Genome Analysers. This can result in a significant reduction in the numbers of clusters that can be analysed. Such low sample diversity is an intrinsic problem of sequencing libraries that are generated by restriction enzyme digestion, such as e4C-seq or reduced-representation libraries. Similarly, this problem can also arise through the combined sequencing of barcoded, multiplexed libraries. We describe a procedure to defer the mapping of cluster coordinates until low-diversity sequences have been passed. This simple procedure can recover substantial amounts of next generation sequencing data that would otherwise be lost.