Correction of sequence-based artifacts in serial analysis of gene expression

Correction of sequence-based artifacts in serial analysis of gene expression
复制标题

DOI:
10.1093/bioinformatics/bth077
复制
发表时间:
2004-05-22
期刊:
影响因子:
5.8
通讯作者:
Wang, CJ
Wang, CJ
中科院分区:
生物学3区
文献类型:
--
作者:
Akmaev, VR;Wang, CJ

文献摘要

被引文献

相似文献

动机:SAGE (Serial Analysis of Gene Expression)是一项通过快速生成大量转录物标签来测量全球基因表达的强大技术。除了在差异基因表达分析中的内在价值外,SAGE标签集合还提供了关于样本转录组大小和形状的丰富信息,并可以加速新基因的发现。这些后一种SAGE应用是由长SAGE的增强方法促进的。基于测序的方法,如SAGE和Long SAGE的一个特点是不可避免地会出现由测序错误引起的伪序列。由于其低随机发生率,这类标签错误对差异表达分析的影响很小。然而,为了充分利用大型SAGE标签数据集的价值,需要考虑和纠正标签工件。结果:我们提出了估计出现的标签错误,并有效的纠错算法。错误率估计是基于随机模型,包括聚合酶链反应和测序错误的贡献。校正算法SAGEScreen是一个多步骤的过程,解决了标记处理、从大量标签中估计经验错误率、相似序列标签分组和观察计数的统计测试。我们将sagesscreen应用于Long SAGE库,并比较几种处理场景的错误率。模拟标签集合的结果表明,sagesscreen纠正了78%的可恢复标签错误,并减少了单标签的出现。
Motivation: Serial Analysis of Gene Expression (SAGE) is a powerful technology for measuring global gene expression, through rapid generation of large numbers of transcript tags. Beyond their intrinsic value in differential gene expression analysis, SAGE tag collections afford abundant information on the size and shape of the sample transcriptome and can accelerate novel gene discovery. These latter SAGE applications are facilitated by the enhanced method of Long SAGE. A characteristic of sequencing-based methods, such as SAGE and Long SAGE is the unavoidable occurrence of artifact sequences resulting from sequencing errors. By virtue of their low-random incidence, such tag errors have minimal impact on differential expression analysis. However, to fully exploit the value of large SAGE tag datasets, it is desirable to account for and correct tag artifacts.Results: We present estimates for occurrences of tag errors, and an efficient error correction algorithm. Error rate estimates are based on a stochastic model that includes the Polymerase chain reaction and sequencing error contributions. The correction algorithm, SAGEScreen, is a multi-step procedure that addresses ditag processing, estimation of empirical error rates from highly abundant tags, grouping of similar-sequence tags and statistical testing of observed counts. We apply SAGEScreen to Long SAGE libraries and compare error rates for several processing scenarios. Results with simulated tag collections indicate that SAGEScreen corrects 78% of recoverable tag errors and reduces the occurrences of singleton tags.