Probabilistic base calling of Solexa sequencing data

Probabilistic base calling of Solexa sequencing data
复制标题

DOI:
10.1186/1471-2105-9-431
复制
发表时间:
2008-10-13
期刊:
影响因子:
3
通讯作者:
Naef, Felix
Naef, Felix
中科院分区:
生物学4区
文献类型:
--
作者:
Rougemont, Jacques;Amzallag, Arnaud;Naef, Felix

文献摘要

被引文献

相似文献

背景资料:Solexa/Illumina短读超高通量DNA测序技术通过DNA菌落的平行合成测序产生数百万个短标签(多达36个碱基)。这样的高通量数据的处理和统计分析提出了新的挑战,目前一个公平的比例的标签被例行丢弃,由于无法将它们匹配到一个参考sequence. Results的,从而降低了有效的吞吐量的technology.Results:我们提出了一种新的基地调用算法,使用基于模型的聚类和概率理论,以确定模糊的基地和代码与IUPAC符号。我们还选择最佳的子标签使用的分数的基础上的信息内容,以消除不确定的碱基对reads.Conclusion的结束:我们表明,该方法提高了基因组的覆盖率和可用的标签数相比,Solexa的数据处理管道平均15%。提供了一个R软件包,可以快速准确地调用Solexa的荧光强度文件,并生成信息丰富的诊断图。
Background: Solexa/Illumina short-read ultra-high throughput DNA sequencing technology produces millions of short tags ( up to 36 bases) by parallel sequencing-by-synthesis of DNA colonies. The processing and statistical analysis of such high-throughput data poses new challenges; currently a fair proportion of the tags are routinely discarded due to an inability to match them to a reference sequence, thereby reducing the effective throughput of the technology.Results: We propose a novel base calling algorithm using model-based clustering and probability theory to identify ambiguous bases and code them with IUPAC symbols. We also select optimal sub-tags using a score based on information content to remove uncertain bases towards the ends of the reads.Conclusion: We show that the method improves genome coverage and number of usable tags as compared with Solexa's data processing pipeline by an average of 15%. An R package is provided which allows fast and accurate base calling of Solexa's fluorescence intensity files and the production of informative diagnostic plots.