Aerolysin Nanopore Identification of Single Nucleotides Using the AdaBoost Model

Aerolysin Nanopore Identification of Single Nucleotides Using the AdaBoost Model
复制标题

DOI:
10.1007/s41664-019-00088-x
复制
发表时间:
2019-04
影响因子:
4.7
通讯作者:
Xuelin Sui;Meng-Yin Li;Yilun Ying;Bingyong Yan;Hui-feng Wang;Jia-Le Zhou;Zhen Gu;Yitao Long
Xuelin Sui;Meng-Yin Li;Yilun Ying;Bingyong Yan;Hui-feng Wang;Jia-Le Zhou;Zhen Gu;Yitao Long
中科院分区:
化学4区
文献类型:
--
作者:
Xuelin Sui;Meng-Yin Li;Yilun Ying;Bingyong Yan;Hui-feng Wang;Jia-Le Zhou;Zhen Gu;Yitao Long

文献摘要

被引文献

相似文献

纳米孔利用单分子堵塞产生的离子电流来识别单分子的结构、构象、化学基团和电荷。尽管在设计敏感成孔材料方面取得了巨大进展,但在某种程度上,具有单组差异的分析物仍然表现出相似的剩余电流或持续时间。剩余电流和持续时间统计结果的严重重叠给混合物中各个单分子的纳米孔区分带来了困难。在本文中,我们提出了基于 AdaBoost 的机器学习模型来识别混合块中具有单组差异的多种分析物。从隐马尔可夫模型 (HMM) 获得的一组特征向量用于训练 AdaBoost 模型。通过采用5′-AAAA-3′(AA3)和5′-GAAA-3′(GA3)的气溶素传感作为模型系统,我们的结果表明AdaBoost模型将识别精度从 ~ 0.293提高到0.991以上。此外,AA3和GA3的五组混合块进一步验证了训练和验证的平均准确率,分别为0.997和0.989。该方法提高了野生型生物纳米孔有效识别单核苷酸差异的能力,无需进行蛋白质设计和实验条件优化。因此,基于AdaBoost的机器学习方法可以促进纳米孔的实际应用,例如遗传和表观遗传检测。
Nanopores employ the ionic current from the single molecule blockage to identify the structure, conformation, chemical groups and charges of a single molecule. Despite the tremendous development in designing sensitive pore-forming materials, at some extent, the analyte with the single group difference still exhibits similar residual current or duration time. The serious overlap in the statistical results of residual current and duration time brings the difficulties in the nanopore discrimination of each single molecules from the mixture. In this paper, we present the AdaBoost-based machine learning model to identify the multiple analyte with single group difference in the mixed blockages. A set of feature vectors which is obtained from Hidden Markov Model (HMM) is used to train the AdaBoost model. By employing the aerolysin sensing of 5ʹ-AAAA-3ʹ (AA3) and 5ʹ-GAAA-3ʹ (GA3) as the model system, our results show that AdaBoost model increases the identification accuracy from ~ 0.293 to above 0.991. Furthermore, five sets of mixed blockages of AA3and GA3further validate the average accuracy of training and validation, which are 0.997 and 0.989, respectively. The proposed methods improve the capacity of wild-type biological nanopore in efficiently identify the single nucleotide difference without designing of protein and optimizing of the experimental condition. Therefore, the AdaBoost-based machine learning approach could promote the nanopore practical application such as genetic and epigenetic detection.