Error rates for nanopore discrimination among cytosine, methylcytosine, and hydroxymethylcytosine along individual DNA strands

Error rates for nanopore discrimination among cytosine, methylcytosine, and hydroxymethylcytosine along individual DNA strands
复制标题

DOI:
10.1073/pnas.1310615110
复制
发表时间:
2013-11-19
影响因子:
11.1
通讯作者:
Akeson, Mark
Akeson, Mark
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Schreiber, Jacob;Wescoe, Zachary L.;Akeson, Mark

文献摘要

被引文献

相似文献

在phi 29 DNA聚合酶的控制下,在单个DNA模板链通过修饰的耻垢分枝杆菌孔蛋白A(M2MspA)纳米孔的易位期间鉴定胞嘧啶、5-甲基胞嘧啶和5-羟甲基胞嘧啶。该鉴定是基于三个连续的离子电流状态,其对应于修饰或未修饰的CG二核苷酸及其紧邻物通过纳米孔限制孔径。为了建立这些调用的质量评分,我们检查了48个不同DNA构建体的类似3,300个易位事件。每个实验分析了带有胞嘧啶、5-甲基胞嘧啶和5-羟甲基胞嘧啶的DNA链的混合物,所述DNA链含有在每个测试分子的靶CG处独立建立正确胞嘧啶甲基化状态的标记物。为了计算这些调用的错误率,我们使用各种机器学习方法建立了决策边界。这些错误率取决于靶向CG二核苷酸的5'和3'端碱基的同一性,并且对于单次读取,范围为1.7%至12.2%。我们估计甲基化状态调用的Q40值(0.01%错误率)可以通过阅读单个分子5-19次来实现,这取决于序列背景。
Cytosine, 5-methylcytosine, and 5-hydroxymethylcytosine were identified during translocation of single DNA template strands through a modified Mycobacterium smegmatis porin A (M2MspA) nanopore under control of phi29 DNA polymerase. This identification was based on three consecutive ionic current states that correspond to passage of modified or unmodified CG dinucleotides and their immediate neighbors through the nanopore limiting aperture. To establish quality scores for these calls, we examined similar to 3,300 translocation events for 48 distinct DNA constructs. Each experiment analyzed a mixture of cytosine-, 5-methylcytosine-, and 5-hydroxymethylcytosine-bearing DNA strands that contained a marker that independently established the correct cytosine methylation status at the target CG of each molecule tested. To calculate error rates for these calls, we established decision boundaries using a variety of machine-learning methods. These error rates depended upon the identity of the bases immediately 5' and 3' of the targeted CG dinucleotide, and ranged from 1.7% to 12.2% for a single-pass read. We estimate that Q40 values (0.01% error rates) for methylation status calls could be achieved by reading single molecules 5-19 times depending upon sequence context.