The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules

The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules
复制标题

DOI:
10.1038/s41597-020-0473-z
复制
发表时间:
2020-05-01
期刊:
影响因子:
9.8
通讯作者:
Tretiak, Sergei
Tretiak, Sergei
中科院分区:
综合性期刊2区
文献类型:
--
作者:
Smith, Justin S.;Zubatyuk, Roman;Tretiak, Sergei

文献摘要

被引文献

相似文献

最大限度地使数据多样化是构建通用和准确的机器学习(ML)模型的中心主题。在化学领域,ML被用来开发预测分子性质的模型,例如量子力学(QM)计算的势能面和原子电荷模型。基于ANI-1x和ANI-1ccx ML的有机分子通用潜力是通过主动学习开发的,这是一个自动化的数据多样化过程。这里,我们描述ANI-1x和ANI-1ccx数据集。为了展示数据的多样性,我们使用降维方案将其可视化,并与现有的数据集进行对比。ANI-1x数据集包含来自5M密度泛函理论计算的多个QM性质,而ANI-1CCX数据集包含通过精确的CCSD(T)/CBS外推获得的500k个数据点。生成该数据大约花费了1400万个CPU核心小时。提供了化学元素C、H、N和O的多种QM计算性质:能量、原子力、多极矩、原子电荷等。我们向社区提供这些数据,以帮助研究和开发化学ML模型。
Maximum diversification of data is a central theme in building generalized and accurate machine learning (ML) models. In chemistry, ML has been used to develop models for predicting molecular properties, for example quantum mechanics (QM) calculated potential energy surfaces and atomic charge models. The ANI-1x and ANI-1ccx ML-based general-purpose potentials for organic molecules were developed through active learning; an automated data diversification process. Here, we describe the ANI-1x and ANI-1ccx data sets. To demonstrate data diversity, we visualize it with a dimensionality reduction scheme, and contrast against existing data sets. The ANI-1x data set contains multiple QM properties from 5 M density functional theory calculations, while the ANI-1ccx data set contains 500 k data points obtained with an accurate CCSD(T)/CBS extrapolation. Approximately 14 million CPU core-hours were expended to generate this data. Multiple QM calculated properties for the chemical elements C, H, N, and O are provided: energies, atomic forces, multipole moments, atomic charges, etc. We provide this data to the community to aid research and development of ML models for chemistry.