MinION Analysis and Reference Consortium: Phase 2 data release and analysis of R9.0 chemistry.

MinION Analysis and Reference Consortium: Phase 2 data release and analysis of R9.0 chemistry.
复制标题

DOI:
10.12688/f1000research.11354.1
复制
发表时间:
2017
期刊:
影响因子:
--
通讯作者:
MinION Analysis and Reference Consortium
MinION Analysis and Reference Consortium
中科院分区:
其他
文献类型:
--
作者:
Jain M;Tyson JR;Loose M;Ip CLC;Eccles DA;O'Grady J;Malla S;Leggett RM;Wallerman O;Jansen HJ;Zalunin V;Birney E;Brown BL;Snutch TP;Olsen HE;MinION Analysis and Reference Consortium

文献摘要

被引文献

相似文献

背景:长读段测序正在迅速发展,并重塑了基因组分析的机会。特别是对于MinION,随着平台和化学的发展,用户群体需要参考数据来设定性能预期并最大限度地利用第三代测序。我们使用R9.0化学对来自大肠杆菌K-12全基因组测序的MinION数据进行了分析,并将结果与旧的R7.3化学进行了比较。 方法:我们计算了MinION读数中插入、缺失和错配的错误率估计值。 结果如下:R9.0的流动池和运行脚本的运行时间特征与针对R7.3化学所观察到的那些相似,但是具有由单个纳米孔处理的每秒碱基的8倍增加(从R7.3和SQK-MAP 005文库制备中的30 bps到R9.0中的250 bps),以及产率随时间的更少下降。二维("2D")N50读取长度与先前的化学反应相比没有变化。使用可定位读段的比例作为碱基识别准确性的量度,来自一维("1D")实验的99.9%的"通过"模板读段是可定位的,并且来自2D实验的约97%。对于1D实验,读段的中位同一性为~89%,对于2D实验,读段的中位同一性为~94%。2D "通过"读段的总错误率(误判+插入+缺失)从R7.3中的9.1%降至R9.0中的7.5%,模板"通过"读段的总错误率从R7.3中的26.7%降至R9.0中的14.5%。 结论:这些2期MinION实验通过提供读取质量、通量和可映射性的估计值作为基线。这些数据集进一步支持开发针对新R9.0化学的生物信息学工具,并为该技术设计新的生物应用。 缩略语:K:千,Kb:酶(一千个碱基对),M:百万,Mb:兆碱基对(一百万个碱基对),Gb:千兆碱基对(十亿个碱基对)。
Background: Long-read sequencing is rapidly evolving and reshaping the suite of opportunities for genomic analysis. For the MinION in particular, as both the platform and chemistry develop, the user community requires reference data to set performance expectations and maximally exploit third-generation sequencing. We performed an analysis of MinION data derived from whole genome sequencing of Escherichia coli K-12 using the R9.0 chemistry, comparing the results with the older R7.3 chemistry. Methods: We computed the error-rate estimates for insertions, deletions, and mismatches in MinION reads. Results: Run-time characteristics of the flow cell and run scripts for R9.0 were similar to those observed for R7.3 chemistry, but with an 8-fold increase in bases per second (from 30 bps in R7.3 and SQK-MAP005 library preparation, to 250 bps in R9.0) processed by individual nanopores, and less drop-off in yield over time. The 2-dimensional (“2D”) N50 read length was unchanged from the prior chemistry. Using the proportion of alignable reads as a measure of base-call accuracy, 99.9% of “pass” template reads from 1-dimensional (“1D”)  experiments were mappable and ~97% from 2D experiments. The median identity of reads was ~89% for 1D and ~94% for 2D experiments. The total error rate (miscall + insertion + deletion ) decreased for 2D “pass” reads from 9.1% in R7.3 to 7.5% in R9.0 and for template “pass” reads from 26.7% in R7.3 to 14.5% in R9.0. Conclusions: These Phase 2 MinION experiments serve as a baseline by providing estimates for read quality, throughput, and mappability. The datasets further enable the development of bioinformatic tools tailored to the new R9.0 chemistry and the design of novel biological applications for this technology. Abbreviations: K: thousand, Kb: kilobase (one thousand base pairs), M: million, Mb: megabase (one million base pairs), Gb: gigabase (one billion base pairs).