Statistical modeling of sequencing errors in SAGE libraries

Statistical modeling of sequencing errors in SAGE libraries
复制标题

DOI:
10.1093/bioinformatics/bth924
复制
发表时间:
2004-08-04
期刊:
影响因子:
5.8
通讯作者:
Speed, Terence P.
Speed, Terence P.
中科院分区:
生物学3区
文献类型:
--
作者:
Beissbarth, Tim;Hyde, Lavinia;Speed, Terence P.

文献摘要

被引文献

相似文献

动机:测序错误可能会使基因表达系列分析(SAGE)所进行的基因表达测量产生偏差。它们可能会引入低丰度的不存在的标签,并降低其他标签的实际丰度。在长SAGE文库中产生的较长标签中,这些影响会加剧。当前的测序技术能够对测序错误率进行相当准确的估计。在此,我们利用SAGE标签的序列邻域以及碱基识别软件的错误估计来校正此类错误。 结果:我们引入了一个用于SAGE中测序错误传播的统计模型,并提出了一种期望最大化(EM)算法,以便根据文库中观察到的序列和碱基识别错误估计来校正这些错误。我们使用模拟的和实验性的SAGE文库对我们的方法进行了测试。在比较SAGE文库时,我们发现测序错误会引入相当大的偏差。高丰度标签可能会被错误地判定为显著差异表达,尤其是在比较具有不同测序错误水平和/或不同大小的文库时。真正差异表达的标签其显著性会降低,因为“真实”标签计数通常被低估。如果接近差异表达阈值的标签被判定为显著,这种情况可能会改变。此外,由于在低丰度时引入了错误标签,文库中存在的不同转录本数量会被高估。我们的校正方法将标签计数调整得更接近真实计数,并且能够部分校正由测序错误引入的偏差。
Motivation: Sequencing errors may bias the gene expression measurements made by Serial Analysis of Gene Expression (SAGE). They may introduce non-existent tags at low abundance and decrease the real abundance of other tags. These effects are increased in the longer tags generated in Long-SAGE libraries. Current sequencing technology generates quite accurate estimates of sequencing error rates. Here we make use of the sequence neighborhood of SAGE tags and error estimates from the base-calling software to correct for such errors.Results: We introduce a statistical model for the propagation of sequencing errors in SAGE and suggest an Expectation-Maximization (EM) algorithm to correct for them given observed sequences in a library and base-calling error estimates. We tested our method using simulated and experimental SAGE libraries. When comparing SAGE libraries, we found that sequencing errors can introduce considerable bias. High abundance tags may be falsely called as significantly differentially expressed, especially when comparing libraries with different levels of sequencing errors and/or of different size. Truly, differentially expressed tags have decreased significance as 'true'-tag counts are generally underestimated. This may alter if tags near the threshold of differential expression are called significant. Moreover, the number of different transcripts present in a library is overestimated as false tags are introduced at low abundance. Our correction method adjusts the tag counts to be closer to the true counts and is able to partly correct for biases introduced by sequencing errors.