Melting point prediction of organic molecules by deciphering the chemical structure into a natural language.

Melting point prediction of organic molecules by deciphering the chemical structure into a natural language.
复制标题

DOI:
10.1039/d0cc07384a
复制
发表时间:
2021-02
影响因子:
4.9
通讯作者:
Weiming Mi;Huijun Chen;Donghua Alan Zhu;Tao Zhang;F. Qian
Weiming Mi;Huijun Chen;Donghua Alan Zhu;Tao Zhang;F. Qian
中科院分区:
化学2区
文献类型:
--
作者:
Weiming Mi;Huijun Chen;Donghua Alan Zhu;Tao Zhang;F. Qian

文献摘要

被引文献

相似文献

在小分子药物的早期发现阶段建立定量的结构-性质关系是非常可取的。利用自然语言处理(NLP),我们提出了一个机器学习模型来处理有机小分子的线形符号,从而能够预测它们的熔点。模型的预测精度得益于对同一分子的不同规范化微笑形式的训练,并且不会随着大小、复杂性和结构灵活性的增加而降低。当使用两种不同的归一化微笑形式组合来训练模型时,预测精度提高。与以往的基于片段或基于描述符的模型不同,这种基于NLP的模型的预测精度不会随着分子的大小、复杂性和结构灵活性的增加而降低。通过将化学结构表示为自然语言,这种基于NLP的模型为药物的发现和开发提供了一种潜在的定量结构-性质预测工具。
Establishing quantitative structure-property relationships for the rational design of small molecule drugs at the early discovery stage is highly desirable. Using natural language processing (NLP), we proposed a machine learning model to process the line notation of small organic molecules, allowing the prediction of their melting points. The model prediction accuracy benefits from training upon different canonicalized SMILES forms of the same molecules and does not decrease with increasing size, complexity, and structural flexibility. When a combination of two different canonicalized SMILES forms is used to train the model, the prediction accuracy improves. Largely distinguished from the previous fragment-based or descriptor-based models, the prediction accuracy of this NLP-based model does not decrease with increasing size, complexity, and structural flexibility of molecules. By representing the chemical structure as a natural language, this NLP-based model offers a potential tool for quantitative structure-property prediction for drug discovery and development.