A deep neural network model for packing density predictions and its application in the study of 1.5 million organic molecules

A deep neural network model for packing density predictions and its application in the study of 1.5 million organic molecules
复制标题

DOI:
10.1039/c9sc02677k
复制
发表时间:
2019-09-28
期刊:
影响因子:
8.4
通讯作者:
Hachmann, Johannes
Hachmann, Johannes
中科院分区:
化学1区
文献类型:
--
作者:
Afzal, Mohammad Atif Faiz;Sonpal, Aditya;Hachmann, Johannes

文献摘要

被引文献

相似文献

开发新化合物和新材料的过程越来越多地受到计算建模和模拟的驱动,这使我们能够在实验室中进行研究之前对候选材料进行表征。有机材料的一个重要特性是它们的体积包装,这高度依赖于它们的分子结构。通过控制后者,我们可以实现具有所需密度(以及其他目标性能)的材料。分子动力学模拟是一种流行且相当准确的计算分子体积密度的方法,然而,由于这些计算是计算密集型的,因此对于大规模评估候选材料的高通量筛选研究来说,它们并不是一种实际可行的选择。在这项工作中,我们使用机器学习来开发数据衍生的预测模型,这是基于物理的模拟的替代方案,我们利用它对150万个小有机分子进行超筛选,并深入了解结构组成和包装密度之间的关系。我们还利用本研究分析了所采用的神经网络方法的学习曲线,并获得了关于模型性能和训练数据大小依赖关系的经验数据,这将为未来的研究提供信息。
The process of developing new compounds and materials is increasingly driven by computational modeling and simulation, which allow us to characterize candidates before pursuing them in the laboratory. One of the non-trivial properties of interest for organic materials is their packing in the bulk, which is highly dependent on their molecular structure. By controlling the latter, we can realize materials with a desired density (as well as other target properties). Molecular dynamics simulations are a popular and reasonably accurate way to compute the bulk density of molecules, however, since these calculations are computationally intensive, they are not a practically viable option for high-throughput screening studies that assess material candidates on a massive scale. In this work, we employ machine learning to develop a data-derived prediction model that is an alternative to physics-based simulations, and we utilize it for the hyperscreening of 1.5 million small organic molecules as well as to gain insights into the relationship between structural makeup and packing density. We also use this study to analyze the learning curve of the employed neural network approach and gain empirical data on the dependence of model performance and training data size, which will inform future investigations.