Analysis of nanopore data using hidden Markov models

Analysis of nanopore data using hidden Markov models
复制标题

DOI:
10.1093/bioinformatics/btv046
复制
发表时间:
2015-06-15
期刊:
影响因子:
5.8
通讯作者:
Karplus, Kevin
Karplus, Kevin
中科院分区:
生物学3区
文献类型:
--
作者:
Schreiber, Jacob;Karplus, Kevin

文献摘要

被引文献

相似文献

动机:基于纳米孔的测序技术可以通过分析生物分子通过孔时产生的依赖于序列的离子电流步骤来重建生物序列的特性。通常,这涉及到将新数据对齐到引用,其中引用构造和对齐都是手工执行的。结果:我们提出了一种自动化方法,通过使用隐马尔可夫模型将纳米孔数据对准参考。从先前的处理步骤和使用的酶类产生的几个特征可以简单地合并到模型中。此前,M2MspA纳米孔被证明具有足够的灵敏度,可以区分胞嘧啶、甲基胞嘧啶和羟甲基胞嘧啶。我们通过自动计算三种胞嘧啶变体之间区分的错误率,在该数据的一个子集上验证了我们的自动化方法,并表明自动化方法产生2-3%的错误率,低于以前手动分割和对齐的10%错误率。
Motivation: Nanopore-based sequencing techniques can reconstruct properties of biosequences by analyzing the sequence-dependent ionic current steps produced as biomolecules pass through a pore. Typically this involves alignment of new data to a reference, where both reference construction and alignment have been performed by hand.Results: We propose an automated method for aligning nanopore data to a reference through the use of hidden Markov models. Several features that arise from prior processing steps and from the class of enzyme used can be simply incorporated into the model. Previously, the M2MspA nanopore was shown to be sensitive enough to distinguish between cytosine, methylcytosine and hydroxymethylcytosine. We validated our automated methodology on a subset of that data by automatically calculating an error rate for the distinction between the three cytosine variants and show that the automated methodology produces a 2-3% error rate, lower than the 10% error rate from previous manual segmentation and alignment..