A novel artificial intelligence-based approach for identification of deoxynucleotide aptamers.

A novel artificial intelligence-based approach for identification of deoxynucleotide aptamers.
复制标题

DOI:
10.1371/journal.pcbi.1009247
复制
发表时间:
2021-08
影响因子:
4.3
通讯作者:
Parés-Matos EI
Parés-Matos EI
中科院分区:
生物学2区
文献类型:
--
作者:
Heredia FL;Roche-Lima A;Parés-Matos EI

文献摘要

参考文献

被引文献

相似文献

通过指数富集配体系统进化(SELEX)方法选择DNA适体涉及多个结合步骤,其中将靶和随机化DNA序列文库混合以选择单个核苷酸特异性分子。通常,完成SELEX需要10到20个步骤。在整个过程中,有必要区分真正的DNA适体和未指定的DNA结合序列。因此,开发了一种新的基于机器学习的方法来支持和简化SELEX过程的早期步骤,以帮助区分DNA适体与DNA结合序列的未指定靶点之间的结合。基于自然语言处理(NLP)和机器学习(ML)实现了识别适体的人工智能(AI)方法。使用NLP方法(CountVectorizer)从核苷酸序列中提取信息。使用来自NLP方法的数据连同序列信息沿着训练四种ML算法(逻辑回归、决策树、高斯朴素贝叶斯、支持向量机)。性能最好的模型是支持向量机,因为它具有区分阳性和阴性类别的最佳能力。在我们的模型中,观察到0.995的准确度(A),模型正确分类的样本比例,以及0.998的接收操作曲线下面积(AUROC),模型能够区分类别的程度。开发的AI方法可用于鉴定潜在的DNA适体以减少SELEX选择中的轮数。这种新的方法可以应用于DNA文库的设计,并导致一个更有效和更快的过程中选择DNA适体在SELEX。在这份手稿中,作者解释了一种新的人工智能方法的开发和验证,以支持和简化SELEX过程的早期步骤,以帮助区分脱氧核苷酸适体与DNA结合序列的未指定目标之间的结合。该方法基于自然语言处理和机器学习实现。使用自然语言处理方法CountVectorizer从核苷酸序列中提取信息。使用来自自然语言处理方法的数据以及序列信息沿着训练四种机器学习算法(逻辑回归、决策树、高斯朴素贝叶斯和支持向量机)。从这四种经过训练的机器学习算法中,最佳性能和所选模型是支持向量机,因为它具有最佳的判别指标(即,准确度(A)= 0.995; AUROC(Au)= 0.998)。一般来说,所有模型都显示出预测DNA适体序列的良好度量结果。机器学习模型的复杂性和难以解释可能会阻碍其应用到标准实践中。为此,已经在开发一个网络应用程序,以促进对所获得结果的解释和应用。
The selection of a DNA aptamer through the Systematic Evolution of Ligands by EXponential enrichment (SELEX) method involves multiple binding steps, in which a target and a library of randomized DNA sequences are mixed for selection of a single, nucleotide-specific molecule. Usually, 10 to 20 steps are required for SELEX to be completed. Throughout this process it is necessary to discriminate between true DNA aptamers and unspecified DNA-binding sequences. Thus, a novel machine learning-based approach was developed to support and simplify the early steps of the SELEX process, to help discriminate binding between DNA aptamers from those unspecified targets of DNA-binding sequences. An Artificial Intelligence (AI) approach to identify aptamers were implemented based on Natural Language Processing (NLP) and Machine Learning (ML). NLP method (CountVectorizer) was used to extract information from the nucleotide sequences. Four ML algorithms (Logistic Regression, Decision Tree, Gaussian Naïve Bayes, Support Vector Machines) were trained using data from the NLP method along with sequence information. The best performing model was Support Vector Machines because it had the best ability to discriminate between positive and negative classes. In our model, an Accuracy (A) of 0.995, the fraction of samples that the model correctly classified, and an Area Under the Receiving Operating Curve (AUROC) of 0.998, the degree by which a model is capable of distinguishing between classes, were observed. The developed AI approach is useful to identify potential DNA aptamers to reduce the amount of rounds in a SELEX selection. This new approach could be applied in the design of DNA libraries and result in a more efficient and faster process for DNA aptamers to be chosen during SELEX. In this manuscript authors explain the development and validation of a novel artificial intelligence approach to support and simplify the early steps of the process from SELEX, to help discriminate binding between deoxynucleotide aptamers from those unspecified targets of DNA-binding sequences. The approach was implemented based on Natural Language Processing and Machine Learning. CountVectorizer, a Natural Language Processing method, was used to extract information from nucleotide sequences. Four Machine Learning algorithms (Logistic Regression, Decision Tree, Gaussian Naïve Bayes, and Support Vector Machines) were trained using data from the Natural Language Processing method along with sequence information. From these four trained machine learning algorithms, the best performance and selected model was Support Vectors Machines, because it had the best discriminatory metrics (i.e., Accuracy (A) = 0.995; AUROC (AU) = 0.998). In general, all models showed good metric results for predicting DNA aptamer sequences. The Machine Learning model complexity and difficult interpretation may hinder its application into the standard practice. For this reason, the development of a web-app is already taking place to facilitate the interpretation and application of the obtained results.
DOI: 10.1038/srep21285
发表时间: 2016-02-22
期刊: Scientific reports
影响因子: 4.6
作者:
Ahirwar R;Nahar S;Aggarwal S;Ramachandran S;Maiti S;Nahar P
通讯作者: Nahar P
DOI: 10.1016/j.talanta.2018.06.035
发表时间: 2018-11-01
期刊: TALANTA
影响因子: 6.1
作者:
Jarczewska, Marta;Rebis, Janusz;Malinowska, Elzbieta
通讯作者: Malinowska, Elzbieta
DOI: 10.3390/ph9020029
发表时间: 2016-05-19
期刊: Pharmaceuticals (Basel, Switzerland)
影响因子: --
作者:
Gijs M;Penner G;Blackler GB;Impens NR;Baatout S;Luxen A;Aerts AM
通讯作者: Aerts AM
DOI: 10.1613/jair.953
发表时间: 2002-01-01
影响因子: 5
作者:
Chawla, NV;Bowyer, KW;Kegelmeyer, WP
通讯作者: Kegelmeyer, WP
DOI: 10.1007/bf00996358
发表时间: 1994-01-01
影响因子: 2.8
作者:
KLUG, SJ;FAMULOK, M
通讯作者: FAMULOK, M