SeqTrim: a high-throughput pipeline for pre-processing any type of sequence read.

SeqTrim: a high-throughput pipeline for pre-processing any type of sequence read.
复制标题

DOI:
10.1186/1471-2105-11-38
复制
发表时间:
2010-01-20
期刊:
影响因子:
3
通讯作者:
Claros MG
Claros MG
中科院分区:
生物学4区
文献类型:
--
作者:
Falgueras J;Lara AJ;Fernández-Pozo N;Cantón FR;Pérez-Trabado G;Claros MG

文献摘要

参考文献

被引文献

相似文献

高通量自动测序使测序数据以指数级的速度增长。这需要提高序列质量和可靠性,以避免数据库受到人工合成序列的污染。焦磷酸测序的到来加强了这一问题,并需要可定制的预处理算法。SeqTrim既作为Web应用程序实现,也作为独立的命令行应用程序实现。已经发表的和新设计的算法已经包括用于识别序列插入,去除低质量、载体、接头、低复杂性和污染序列,以及检测嵌合阅读。由于提供了多种输入和输出格式,因此可以将其包含在序列处理工作流中。由于其特定的算法,SeqTrim的性能优于作为Web服务或独立应用程序实现的其他预处理器。它与来自EST文库、SSH文库、基因组DNA文库和焦磷酸测序读数的序列同样表现良好,并且不会导致过度修剪。SeqTrim是一种高效的流水线,设计用于任何类型的序列读取的预处理,包括下一代测序。它很容易配置,并提供了一个友好的界面,允许用户了解每个预处理阶段的序列发生了什么,如果需要的话,还可以验证单个序列的预处理。与先前描述的预处理器相比,推荐的流水线揭示了关于每个序列的更多信息,并且可以丢弃更多的测序或实验人工制品。
High-throughput automated sequencing has enabled an exponential growth rate of sequencing data. This requires increasing sequence quality and reliability in order to avoid database contamination with artefactual sequences. The arrival of pyrosequencing enhances this problem and necessitates customisable pre-processing algorithms. SeqTrim has been implemented both as a Web and as a standalone command line application. Already-published and newly-designed algorithms have been included to identify sequence inserts, to remove low quality, vector, adaptor, low complexity and contaminant sequences, and to detect chimeric reads. The availability of several input and output formats allows its inclusion in sequence processing workflows. Due to its specific algorithms, SeqTrim outperforms other pre-processors implemented as Web services or standalone applications. It performs equally well with sequences from EST libraries, SSH libraries, genomic DNA libraries and pyrosequencing reads and does not lead to over-trimming. SeqTrim is an efficient pipeline designed for pre-processing of any type of sequence read, including next-generation sequencing. It is easily configurable and provides a friendly interface that allows users to know what happened with sequences at every pre-processing stage, and to verify pre-processing of an individual sequence if desired. The recommended pipeline reveals more information about each sequence than previously described pre-processors and can discard more sequencing or experimental artefacts.
DOI: 10.1093/nar/gkm369
发表时间: 2007-07
影响因子: 14.9
作者:
Lee B;Hong T;Byun SJ;Woo T;Choi YJ
通讯作者: Choi YJ
DOI: 10.1093/nar/gkm378
发表时间: 2007-07
影响因子: 14.9
作者:
Nagaraj SH;Deshpande N;Gasser RB;Ranganathan S
通讯作者: Ranganathan S
DOI: 10.1093/bioinformatics/btm632
发表时间: 2008-02-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
White, James Robert;Roberts, Michael;Pop, Mihai
通讯作者: Pop, Mihai
DOI: 10.1093/bioinformatics/15.2.106
发表时间: 1999-02-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Seluja, GA;Farmer, A;Schad, PA
通讯作者: Schad, PA
DOI: 10.1186/1471-2105-9-5
发表时间: 2008-01-07
期刊: BMC bioinformatics
影响因子: 3
作者:
Forment J;Gilabert F;Robles A;Conejero V;Nuez F;Blanca JM
通讯作者: Blanca JM