Quality filtering of Illumina index reads mitigates sample cross-talk.

Quality filtering of Illumina index reads mitigates sample cross-talk.
复制标题

DOI:
10.1186/s12864-016-3217-x
复制
发表时间:
2016-11-04
期刊:
影响因子:
4.4
通讯作者:
Vetsigian KH
Vetsigian KH
中科院分区:
生物学2区
文献类型:
--
作者:
Wright ES;Vetsigian KH

文献摘要

参考文献

被引文献

相似文献

在 Illumina 测序过程中对多个样本进行多重分析是一种常见做法,并且随着平台吞吐量的增加,其重要性也迅速增长。多路分离过程中的错误分配(其中序列与错误的样本相关)是 Illumina 测序平台上被忽视的错误模式。这导致多重样本之间的串扰率较低,并且可能在需要检测罕见变异或多重样本时产生不利影响。当使用独特的 i5 和 i7 索引序列复用 14 个不同的样本时,我们观察到串扰率平均为 0.24%。该串扰率相当于 Illumina HiSeq 2500 单个泳道上的 254,632 个错误分配读数。值得注意的是,所有类型的错误分配发生率相似:不正确的 i5、不正确的 i7 和不正确的序列读数。我们证明,通过索引读数的质量过滤几乎可以消除错误分配,同时保留约 90% 的原始序列。多重样本之间的串扰是 Illumina 平台上的一个重要错误模式,特别是当样本仅由单个唯一索引分隔时。索引序列的质量过滤为最大限度地减少样本之间的串扰提供了有效的解决方案。此外,我们提出了一种简单的方法来验证样本之间的串扰程度并优化质量分数阈值,该方法不需要额外的控制样本,甚至可以在之前的运行中事后执行。本文的在线版本 (doi:10.1186/s12864-016-3217-x) 包含补充材料,可供授权用户使用。
Multiplexing multiple samples during Illumina sequencing is a common practice and is rapidly growing in importance as the throughput of the platform increases. Misassignments during de-multiplexing, where sequences are associated with the wrong sample, are an overlooked error mode on the Illumina sequencing platform. This results in a low rate of cross-talk among multiplexed samples and can cause detrimental effects in studies requiring the detection of rare variants or when multiplexing a large number of samples. We observed rates of cross-talk averaging 0.24 % when multiplexing 14 different samples with unique i5 and i7 index sequences. This cross-talk rate corresponded to 254,632 misassigned reads on a single lane of the Illumina HiSeq 2500. Notably, all types of misassignment occur at similar rates: incorrect i5, incorrect i7, and incorrect sequence reads. We demonstrate that misassignments can be nearly eliminated by quality filtering of index reads while preserving about 90 % of the original sequences. Cross-talk among multiplexed samples is a significant error mode on the Illumina platform, especially if samples are only separated by a single unique index. Quality filtering of index sequences offers an effective solution to minimizing cross-talk among samples. Furthermore, we propose a straightforward method for verifying the extent of cross-talk between samples and optimizing quality score thresholds that does not require additional control samples and can even be performed post hoc on previous runs. The online version of this article (doi:10.1186/s12864-016-3217-x) contains supplementary material, which is available to authorized users.
DOI: 10.1093/nar/gkr344
发表时间: 2011-07
影响因子: 14.9
作者:
Nakamura K;Oshima T;Morimoto T;Ikeda S;Yoshikawa H;Shiwa Y;Ishikawa S;Linak MC;Hirai A;Takahashi H;Altaf-Ul-Amin M;Ogasawara N;Kanaya S
通讯作者: Kanaya S
DOI: 10.1093/nar/gku1341
发表时间: 2015-03-31
影响因子: 14.9
作者:
Schirmer M;Ijaz UZ;D'Amore R;Hall N;Sloan WT;Quince C
通讯作者: Quince C
DOI: 10.1186/1471-2105-11-485
发表时间: 2010-09-27
期刊: BMC bioinformatics
影响因子: 3
作者:
Cox MP;Peterson DA;Biggs PJ
通讯作者: Biggs PJ
DOI: 10.1371/journal.pone.0094249
发表时间: 2014
期刊: PloS one
影响因子: 3.7
作者:
Nelson MC;Morrison HG;Benjamino J;Grim SL;Graf J
通讯作者: Graf J
DOI: 10.1093/nar/gkq1019
发表时间: 2011-01
影响因子: 14.9
作者:
Leinonen R;Sugawara H;Shumway M;International Nucleotide Sequence Database Collaboration
通讯作者: International Nucleotide Sequence Database Collaboration