ConDeTri--a content dependent read trimmer for Illumina data.

ConDeTri--a content dependent read trimmer for Illumina data.
复制标题

DOI:
10.1371/journal.pone.0026314
复制
发表时间:
2011
期刊:
影响因子:
3.7
通讯作者:
Künstner A
Künstner A
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Smeds L;Künstner A

文献摘要

参考文献

被引文献

相似文献

在过去的几年中,DNA和RNA测序已经开始在生物和医学应用中发挥越来越重要的作用,特别是由于新测序仪产生的测序数据量更大以及测序成本的大幅降低。特别是,Illumina/Solexa测序对从模型和非模型生物体收集数据产生了越来越大的影响。然而,尚未建立准确和易于使用的质量过滤工具。我们提出了ConDeTri,一种使用每个单独碱基的质量分数对下一代测序数据进行内容依赖性读段修剪的方法。该方法的主要重点是从读数中去除测序错误,以便可以标准化测序读数。该方法的另一方面是在下一代测序数据处理和分析流水线中并入读段修剪。它可以处理任意长度的单端和双端序列数据,并且独立于测序覆盖率和用户交互。ConDeTri能够修剪和删除低质量分数的读数,以节省从头组装过程中的计算时间和内存使用。低覆盖率或大型基因组测序项目将特别受益于修剪读数。该方法可以很容易地集成到Illumina数据的预处理和分析管道中。可在http://code.google.com/p/condetri网站上免费获得。
During the last few years, DNA and RNA sequencing have started to play an increasingly important role in biological and medical applications, especially due to the greater amount of sequencing data yielded from the new sequencing machines and the enormous decrease in sequencing costs. Particularly, Illumina/Solexa sequencing has had an increasing impact on gathering data from model and non-model organisms. However, accurate and easy to use tools for quality filtering have not yet been established. We present ConDeTri, a method for content dependent read trimming for next generation sequencing data using quality scores of each individual base. The main focus of the method is to remove sequencing errors from reads so that sequencing reads can be standardized. Another aspect of the method is to incorporate read trimming in next-generation sequencing data processing and analysis pipelines. It can process single-end and paired-end sequence data of arbitrary length and it is independent from sequencing coverage and user interaction. ConDeTri is able to trim and remove reads with low quality scores to save computational time and memory usage during de novo assemblies. Low coverage or large genome sequencing projects will especially gain from trimming reads. The method can easily be incorporated into preprocessing and analysis pipelines for Illumina data. Freely available on the web at http://code.google.com/p/condetri.
DOI: 10.1186/gb-2011-12-3-r31
发表时间: 2011
期刊: Genome biology
影响因子: 12.3
作者:
Ye L;Hillier LW;Minx P;Thane N;Locke DP;Martin JC;Chen L;Mitreva M;Miller JR;Haub KV;Dooling DJ;Mardis ER;Wilson RK;Weinstock GM;Warren WC
通讯作者: Warren WC
DOI: 10.1101/gr.097261.109
发表时间: 2010-02-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Li, Ruiqiang;Zhu, Hongmei;Wang, Jun
通讯作者: Wang, Jun
DOI: 10.1186/1471-2105-11-130
发表时间: 2010-03-15
期刊: BMC bioinformatics
影响因子: 3
作者:
Ratan A;Zhang Y;Hayes VM;Schuster SC;Miller W
通讯作者: Miller W
DOI: 10.1093/bioinformatics/btq653
发表时间: 2011-02-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Ilie, Lucian;Fazayeli, Farideh;Ilie, Silvana
通讯作者: Ilie, Silvana
DOI: 10.1186/gb-2010-11-11-r116
发表时间: 2010
期刊: Genome biology
影响因子: 12.3
作者:
Kelley DR;Schatz MC;Salzberg SL
通讯作者: Salzberg SL