Shining a light on dark sequencing: characterising errors in Ion Torrent PGM data.

Shining a light on dark sequencing: characterising errors in Ion Torrent PGM data.
复制标题

DOI:
10.1371/journal.pcbi.1003031
复制
发表时间:
2013-04
影响因子:
4.3
通讯作者:
Tyson GW
Tyson GW
中科院分区:
生物学2区
文献类型:
--
作者:
Bragg LM;Stone G;Butler MK;Hugenholtz P;Tyson GW

文献摘要

参考文献

被引文献

相似文献

Ion Torrent个人基因组机(PGM)是一种新的测序平台,与其他测序技术有很大不同,它通过测量pH值而不是光来检测聚合事件。使用重测序数据集,我们全面描述了PGM在基础和流程水平上引入的偏差和误差,包括芯片密度、测序试剂盒、模板种类和机器等因素的组合。我们发现两种不同的插入/删除(indel)错误类型占了PGM引入的大多数错误。主要误差来源是不准确的流调用,使用OneTouch 200 bp试剂盒以2.84%的原始率(质量裁剪后为1.38%)引入索引。不准确的流量调用通常会导致所谓的短均聚物和所谓的长均聚物。流量调用精度随着连续的流量循环而下降,但我们也发现流量错误率有显著的周期性波动,对应于流量循环模式中的特定位置。另一种不太常见的PGM错误,高频索引(HFI)错误,是相对于参考基因组中给定的碱基位置,在读取中以非常高的频率出现的索引,但在大多数情况下,在不同的运行中不能一致地复制。HFI误差在参考文献中大约每一千个碱基中出现一次,对应于reads中0.06%的碱基。目前,PGM还没有达到与之竞争的光基技术的精度。然而,流量调用的不准确性是系统性的,这里开发的流量值统计模型将使pgm特定的生物信息学方法得以开发,这将解释这些错误。HFI错误可能更具挑战性,特别是对于多态性和扩增子应用,但可以通过跨多个芯片对相同的DNA模板进行测序来克服。在生物学中,DNA测序通常用于揭示活生物体的遗传信息。近年来,技术进步带来了高通量、低成本的DNA测序仪(“测序仪”)。2011年,生命科学公司发布了一款新的测序仪——离子激流个人基因组机(PGM)。这是第一个测量pH值变化的测序仪,而不是发射光来记录测序反应。因此,这种独特的技术既具有成本效益,又具有很高的准确性,使其对许多实验室具有吸引力。然而,每一种测序技术都会在得到的DNA序列中引入独特的误差和偏差,了解pgm特异性特征对于确定这种新技术的合适应用至关重要。我们全面检查了pgm测序数据的错误和偏差类型,包括芯片密度、模板试剂盒、模板DNA和两台机器。利用统计方法,我们量化了实验变量的影响,以及DNA序列特异性效应,并发现PGM有两种类型的技术特异性误差。我们还发现PGM的精度比基于光的技术差,我们对该技术提出了建议,并提供了克服PGM测序误差的统计模型。
The Ion Torrent Personal Genome Machine (PGM) is a new sequencing platform that substantially differs from other sequencing technologies by measuring pH rather than light to detect polymerisation events. Using re-sequencing datasets, we comprehensively characterise the biases and errors introduced by the PGM at both the base and flow level, across a combination of factors, including chip density, sequencing kit, template species and machine. We found two distinct insertion/deletion (indel) error types that accounted for the majority of errors introduced by the PGM. The main error source was inaccurate flow-calls, which introduced indels at a raw rate of 2.84% (1.38% after quality clipping) using the OneTouch 200 bp kit. Inaccurate flow-calls typically resulted in over-called short-homopolymers and under-called long-homopolymers. Flow-call accuracy decreased with consecutive flow cycles, but we also found significant periodic fluctuations in the flow error-rate, corresponding to specific positions within the flow-cycle pattern. Another less common PGM error, high frequency indel (HFI) errors, are indels that occur at very high frequency in the reads relative to a given base position in the reference genome, but in the majority of instances were not replicated consistently across separate runs. HFI errors occur approximately once every thousand bases in the reference, and correspond to 0.06% of bases in reads. Currently, the PGM does not achieve the accuracy of competing light-based technologies. However, flow-call inaccuracy is systematic and the statistical models of flow-values developed here will enable PGM-specific bioinformatics approaches to be developed, which will account for these errors. HFI errors may prove more challenging to address, especially for polymorphism and amplicon applications, but may be overcome by sequencing the same DNA template across multiple chips. DNA sequencing is used routinely within biology to reveal the genetic information of living organisms. In recent years, technological advances have led to the availability of high-throughput, low-cost DNA sequencing machines (‘sequencers’). In 2011, Life Sciences released a new sequencer, the Ion Torrent Personal Genome Machine (PGM). This is the first sequencer to measure changes in pH rather that emitted light to register sequencing reactions. Consequently, this unique technology is both cost-effective and advertised to have high accuracy, making it attractive for many laboratories. However, every sequencing technology introduces unique errors and biases into the resulting DNA sequences, and understanding PGM-specific characteristics is crucial to determining suitable applications for this new technology. We comprehensively examine the types of errors and biases in PGM-sequenced data across several experimental variables, including chip density, template kit, template DNA and across two machines. Using statistical approaches, we quantify the influence of experimental variables, as well as DNA sequence-specific effects, and find that the PGM has two types of technology-specific errors. We also find that the accuracy of the PGM is poorer than that of light-based technologies, and we make recommendations for this technology as well as provide statistical models for overcoming PGM sequencing errors.
DOI: 10.1186/gb-2007-8-7-r143
发表时间: 2007
期刊: Genome biology
影响因子: 12.3
作者:
Huse SM;Huber JA;Morrison HG;Sogin ML;Welch DM
通讯作者: Welch DM
DOI: 10.1093/bioinformatics/btq365
发表时间: 2010-09-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Balzer S;Malde K;Lanzén A;Sharma A;Jonassen I
通讯作者: Jonassen I
DOI: 10.1038/nbt.2198
发表时间: 2012-05-01
影响因子: 46.9
作者:
Loman, Nicholas J.;Misra, Raju V.;Pallen, Mark J.
通讯作者: Pallen, Mark J.
DOI: 10.1038/nmeth.1361
发表时间: 2009-09-01
期刊: NATURE METHODS
影响因子: 48
作者:
Quince, Christopher;Lanzen, Anders;Sloan, William T.
通讯作者: Sloan, William T.
DOI: 10.1186/1756-0500-4-149
发表时间: 2011-05-26
期刊: BMC research notes
影响因子: 1.8
作者:
Jérôme M;Noirot C;Klopp C
通讯作者: Klopp C