Controlled exploration of chemical space by machine learning of coarse-grained representations.

Controlled exploration of chemical space by machine learning of coarse-grained representations.
复制标题

DOI:
10.1103/physreve.100.033302
复制
发表时间:
2019-05
期刊:
Physical review. E
影响因子:
--
通讯作者:
Christian Hoffmann;R. Menichetti;K. Kanekal;T. Bereau
Christian Hoffmann;R. Menichetti;K. Kanekal;T. Bereau
中科院分区:
其他
文献类型:
--
作者:
Christian Hoffmann;R. Menichetti;K. Kanekal;T. Bereau

文献摘要

被引文献

相似文献

化合物空间太大,无法穷尽探测。这导致高通量协议大幅子采样,并导致稀疏和不均匀的数据集。我们不是任意选择化合物,而是根据感兴趣的目标性质系统地探索化学空间。我们首先通过引入马尔可夫链蒙特卡罗方案跨化合物进行重要性抽样。然后,我们在采样数据上训练机器学习(ML)模型,以扩展探测的化学空间区域。我们的提升过程将化合物的数量提高了2到10倍,这是由ML模型的粗粒度表示实现的,这既简化了结构-性质关系,又减小了化学空间的大小。ML模型正确地恢复了转移自由能之间的线性关系。这些线性关系对应于数据集的全局特征,标记了预测可靠的化学空间区域;这是预测方差的更鲁棒的替代方案。将粗粒度模拟与ML相结合,产生了一个前所未有的130万种化合物的药物膜插入自由能数据库。
The size of chemical compound space is too large to be probed exhaustively. This leads high-throughput protocols to drastically subsample and results in sparse and nonuniform datasets. Rather than arbitrarily selecting compounds, we systematically explore chemical space according to the target property of interest. We first perform importance sampling by introducing a Markov chain Monte Carlo scheme across compounds. We then train a machine learning (ML) model on the sampled data to expand the region of chemical space probed. Our boosting procedure enhances the number of compounds by a factor 2 to 10, enabled by the ML model's coarse-grained representation, which both simplifies the structure-property relationship and reduces the size of chemical space. The ML model correctly recovers linear relationships between transfer free energies. These linear relationships correspond to features that are global to the dataset, marking the region of chemical space up to which predictions are reliable; this is a more robust alternative to the predictive variance. Bridging coarse-grained simulations with ML gives rise to an unprecedented database of drug-membrane insertion free energies for 1.3 million compounds.