Sequencing DNA with nanopores: Troubles and biases.

Sequencing DNA with nanopores: Troubles and biases.
复制标题

DOI:
10.1371/journal.pone.0257521
复制
发表时间:
2021
期刊:
影响因子:
3.7
通讯作者:
Nicolas J
Nicolas J
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Delahaye C;Nicolas J

文献摘要

参考文献

被引文献

相似文献

Oxford Nanopore Technologies(ONT)长读测序仪提供了比前几代测序仪更长的DNA片段,但错误率更高。虽然许多论文已经研究了读取校正方法,但很少有人讨论了观察到的错误的详细特征,这是一项因ONT技术中化学和软件的频繁变化而复杂化的任务。MinION测序仪现在更加稳定,本文使用最成熟的流动池和碱基识别器对其错误情况进行了最新的观察。我们研究了细菌和人类DNA读数的纳米孔测序误差偏差。我们发现,尽管纳米孔测序预计不会受到GC偏差的影响,但它是错误的关键参数。特别是,低GC读数比高GC读数具有更少的错误(分别约6%和8%)。均聚区域或具有短重复序列的区域的错误分布(约一半的测序错误的来源)也取决于GC率,并且主要显示缺失,尽管存在一些具有长插入的读段。另一个有趣的发现是,质量度量虽然被高估了,但它提供了有价值的信息来预测错误率和读取的丰度。我们用油菜籽RNA读段集的分析补充了这项研究,并在这些数据中显示了较高水平的错误和较高水平的缺失。最后,我们实现了一个开源管道,用于长期监测误差分布,使用户能够轻松计算本工作中提出的各种分析,包括测序设备的未来开发。总的来说,我们希望这项工作将提供一个更好的纠错方法的设计基础。
Oxford Nanopore Technologies’ (ONT) long read sequencers offer access to longer DNA fragments than previous sequencer generations, at the cost of a higher error rate. While many papers have studied read correction methods, few have addressed the detailed characterization of observed errors, a task complicated by frequent changes in chemistry and software in ONT technology. The MinION sequencer is now more stable and this paper proposes an up-to-date view of its error landscape, using the most mature flowcell and basecaller. We studied Nanopore sequencing error biases on both bacterial and human DNA reads. We found that, although Nanopore sequencing is expected not to suffer from GC bias, it is a crucial parameter with respect to errors. In particular, low-GC reads have fewer errors than high-GC reads (about 6% and 8% respectively). The error profile for homopolymeric regions or regions with short repeats, the source of about half of all sequencing errors, also depends on the GC rate and mainly shows deletions, although there are some reads with long insertions. Another interesting finding is that the quality measure, although over-estimated, offers valuable information to predict the error rate as well as the abundance of reads. We supplemented this study with an analysis of a rapeseed RNA read set and shown a higher level of errors with a higher level of deletion in these data. Finally, we have implemented an open source pipeline for long-term monitoring of the error profile, which enables users to easily compute various analysis presented in this work, including for future developments of the sequencing device. Overall, we hope this work will provide a basis for the design of better error-correction methods.
DOI: 10.1186/s13059-018-1462-9
发表时间: 2018-07-13
期刊: Genome biology
影响因子: 12.3
作者:
Rang FJ;Kloosterman WP;de Ridder J
通讯作者: de Ridder J
DOI: 10.1093/nar/gkr344
发表时间: 2011-07
影响因子: 14.9
作者:
Nakamura K;Oshima T;Morimoto T;Ikeda S;Yoshikawa H;Shiwa Y;Ishikawa S;Linak MC;Hirai A;Takahashi H;Altaf-Ul-Amin M;Ogasawara N;Kanaya S
通讯作者: Kanaya S
DOI: 10.1093/nargab/lqz015
发表时间: 2020-03
影响因子: 4.6
作者:
Marchet C;Morisse P;Lecompte L;Lefebvre A;Lecroq T;Peterlongo P;Limasset A
通讯作者: Limasset A
DOI: 10.1039/c5mb00750j
发表时间: 2016-01-01
影响因子: --
作者:
Shin, Sunguk;Park, Joonhong
通讯作者: Park, Joonhong
DOI: 10.1093/bioinformatics/btv540
发表时间: 2016-01-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Leggett RM;Heavens D;Caccamo M;Clark MD;Davey RP
通讯作者: Davey RP