pRESTO: a toolkit for processing high-throughput sequencing raw reads of lymphocyte receptor repertoires

pRESTO: a toolkit for processing high-throughput sequencing raw reads of lymphocyte receptor repertoires
复制标题

DOI:
10.1093/bioinformatics/btu138
复制
发表时间:
2014-07-01
期刊:
影响因子:
5.8
通讯作者:
Kleinstein, Steven H.
Kleinstein, Steven H.
中科院分区:
生物学3区
文献类型:
--
作者:
Vander Heiden, Jason A.;Yaari, Gur;Kleinstein, Steven H.

文献摘要

被引文献

相似文献

在技术进步的推动下,通过高通量测序大规模表征淋巴细胞受体库现在是可行的。虽然有希望,但高种系和体细胞多样性,特别是B细胞免疫球蛋白库,对需要开发专门计算管道的分析提出了挑战。我们开发了REpertoire Sequencing TOolkit(pRESTO),用于处理来自高通量淋巴细胞受体研究的读数。pRESTO处理原始序列以产生纠错的、排序的和注释的序列集,沿着每个步骤的大量度量。该工具包支持多重引物池、单端或双端读取以及使用单分子标识符的新兴技术。pRESTO已在Roche和Illumina平台生成的数据上进行了测试。它具有在可用处理器之间并行化工作的内置能力,并且能够有效地处理由典型的高吞吐量项目生成的数百万个序列。可用性和实施:pRESTO可免费供学术使用。软件包和详细的教程可以从http://clip.med.yale.edu/presto下载。
Driven by dramatic technological improvements, large-scale characterization of lymphocyte receptor repertoires via high-throughput sequencing is now feasible. Although promising, the high germline and somatic diversity, especially of B-cell immunoglobulin repertoires, presents challenges for analysis requiring the development of specialized computational pipelines. We developed the REpertoire Sequencing TOolkit (pRESTO) for processing reads from high-throughput lymphocyte receptor studies. pRESTO processes raw sequences to produce error-corrected, sorted and annotated sequence sets, along with a wealth of metrics at each step. The toolkit supports multiplexed primer pools, single-or paired-end reads and emerging technologies that use single-molecule identifiers. pRESTO has been tested on data generated from Roche and Illumina platforms. It has a built-in capacity to parallelize the work between available processors and is able to efficiently process millions of sequences generated by typical high-throughput projects. Availability and implementation: pRESTO is freely available for academic use. The software package and detailed tutorials may be downloaded from http://clip.med.yale.edu/presto.