Identification of errors introduced during high throughput sequencing of the T cell receptor repertoire.

Identification of errors introduced during high throughput sequencing of the T cell receptor repertoire.
复制标题

DOI:
10.1186/1471-2164-12-106
复制
发表时间:
2011-02-11
期刊:
影响因子:
4.4
通讯作者:
Geiger TL
Geiger TL
中科院分区:
生物学2区
文献类型:
--
作者:
Nguyen P;Ma J;Pei D;Obert C;Cheng C;Geiger TL

文献摘要

被引文献

相似文献

大规模平行测序的最新进展已经将T细胞受体(TCR)库可以探测的深度增加了> 3log 10,从而允许免疫库的饱和测序。这种测序的分辨率取决于其准确性,并且对高通量库分析期间形成的错误的直接评估是有限的。我们使用测序分析了来自TCR转基因Rag-/-小鼠的3种单克隆TCR。使用三分叉设计对每个TCR进行总共27个测序反应,其中样品在重要的处理接合点处被分成3份。分析了超过2000万个互补决定区(CDR)3序列。对低质量序列的过滤减少了但没有消除序列错误,其发生在1-6%的序列内。错误的序列主要是正确的长度,并含有单核苷酸取代。特定的取代率变化显着的位置依赖性的方式。四个取代,所有嘌呤嘧啶颠换,占主导地位。固相扩增和测序而不是液体样品扩增和制备似乎是误差的主要来源。多克隆库的分析证明了误差累积对数据参数的影响。由于错误序列读数的潜在污染,在解释库数据时需要谨慎。然而,错误与phred得分的高度关联、错误序列与亲本序列的高度相关性、特定nt取代的优势以及错误序列中正向与反向读段的偏斜比率指示从库数据集中过滤错误序列的方法。
Recent advances in massively parallel sequencing have increased the depth at which T cell receptor (TCR) repertoires can be probed by >3log10, allowing for saturation sequencing of immune repertoires. The resolution of this sequencing is dependent on its accuracy, and direct assessments of the errors formed during high throughput repertoire analyses are limited. We analyzed 3 monoclonal TCR from TCR transgenic, Rag-/- mice using Illumina® sequencing. A total of 27 sequencing reactions were performed for each TCR using a trifurcating design in which samples were divided into 3 at significant processing junctures. More than 20 million complementarity determining region (CDR) 3 sequences were analyzed. Filtering for lower quality sequences diminished but did not eliminate sequence errors, which occurred within 1-6% of sequences. Erroneous sequences were pre-dominantly of correct length and contained single nucleotide substitutions. Rates of specific substitutions varied dramatically in a position-dependent manner. Four substitutions, all purine-pyrimidine transversions, predominated. Solid phase amplification and sequencing rather than liquid sample amplification and preparation appeared to be the primary sources of error. Analysis of polyclonal repertoires demonstrated the impact of error accumulation on data parameters. Caution is needed in interpreting repertoire data due to potential contamination with mis-sequence reads. However, a high association of errors with phred score, high relatedness of erroneous sequences with the parental sequence, dominance of specific nt substitutions, and skewed ratio of forward to reverse reads among erroneous sequences indicate approaches to filter erroneous sequences from repertoire data sets.