SMILES Pair Encoding: A Data-Driven Substructure Tokenization Algorithm for Deep Learning

SMILES Pair Encoding: A Data-Driven Substructure Tokenization Algorithm for Deep Learning
复制标题

DOI:
10.1021/acs.jcim.0c01127
复制
发表时间:
2021-03-15
影响因子:
5.6
通讯作者:
Fourches, Denis
Fourches, Denis
中科院分区:
化学2区
文献类型:
--
作者:
Li, Xinhao;Fourches, Denis

文献摘要

被引文献

相似文献

基于简化分子输入行输入系统(SMILES)的深度学习模型正逐渐成为化学信息学的一个重要研究课题。在这项研究中,我们介绍了SMILES对编码(SPE),数据驱动的标记化算法。SPE首先从大型化学数据集学习高频SMILES子串的词汇表(例如,ChEMBL),然后根据学习的词汇对SMILES进行标记,用于深度学习模型的实际训练。SPE通过添加人类可读和化学上可解释的SMILES子字符串作为标记,增强了广泛使用的原子级标记化。实例研究表明,SPE在分子生成和定量构效关系(QSAR)预测任务上都能取得上级性能。特别是,基于SPE的生成模型在新奇,多样性和模拟训练集分布的能力方面优于原子级标记化模型。使用24个基准数据集评估了基于SPE的QSAR预测模型的性能,其中SPE始终匹配或优于原子级和k-mer标记化。因此,SPE可能是基于SMILES的深度学习模型的一种有前途的标记化方法。开发了一个开源Python包SmilesPE来实现这个算法,现在可以在https://gihub.com/XinhaoLi74/SmilesPE上免费获得。
Simplified molecular input line entry system (SMILES)-based deep learning models are slowly emerging as an important research topic in cheminformatics. In this study, we introduce SMILES pair encoding (SPE), a data-driven tokenization algorithm. SPE first learns a vocabulary of high-frequency SMILES substrings from a large chemical dataset (e.g., ChEMBL) and then tokenizes SMILES based on the learned vocabulary for the actual training of deep learning models. SPE augments the widely used atom-level tokenization by adding human-readable and chemically explainable SMILES substrings as tokens. Case studies show that SPE can achieve superior performances on both molecular generation and quantitative structure-activity relationship (QSAR) prediction tasks. In particular, the SPE-based generative models outperformed the atom-level tokenization model in the aspects of novelty, diversity, and ability to resemble the training set distribution. The performance of SPE-based QSAR prediction models were evaluated using 24 benchmark datasets where SPE consistently either did match or outperform atom-level and k-mer tokenization. Therefore, SPE could be a promising tokenization method for SMILES-based deep learning models. An open-source Python package SmilesPE was developed to implement this algorithm and is now freely available at https://gihub.com/XinhaoLi74/SmilesPE.