Comparing molecules and solids across structural and alchemical space

Comparing molecules and solids across structural and alchemical space
复制标题

DOI:
10.1039/c6cp00415f
复制
发表时间:
2016-05-28
影响因子:
3.3
通讯作者:
Ceriotti, Michele
Ceriotti, Michele
中科院分区:
化学2区
文献类型:
--
作者:
De, Sandip;Bartok, Albert P.;Ceriotti, Michele

文献摘要

被引文献

相似文献

在开发自动导航复杂材料构型空间的算法时,评估结晶、无序和分子化合物的(散乱)相似性是关键的一步。例如,结构相似性度量对于对结构进行分类、在化学空间中搜索更好的化合物和材料以及推动下一代机器学习技术来预测分子和材料的稳定性和性质至关重要。在过去的几年里,已经设计了几种策略来比较原子配位环境。特别是,原子位置的平滑重叠(SOAP)已经成为获得原子组的平移、旋转和排列不变描述符的一个优雅的框架,这是各种类型的机器学习的原子间势发展的基础。在这里,我们讨论如何使用正则化的熵匹配(重新匹配)方法来组合这样的局部描述符来描述整个分子结构和整体周期结构的相似性,引入了强大的度量,使得能够在统一的框架内导航炼金术和结构的复杂性。此外,使用这个核和岭回归方法,我们可以预测一个有机小分子数据库的原子化能,平均绝对误差低于1kcal mol(-1),这是机器学习技术应用于分子性质评估的一个重要里程碑。
Evaluating the (dis)similarity of crystalline, disordered and molecular compounds is a critical step in the development of algorithms to navigate automatically the configuration space of complex materials. For instance, a structural similarity metric is crucial for classifying structures, searching chemical space for better compounds and materials, and driving the next generation of machine-learning techniques for predicting the stability and properties of molecules and materials. In the last few years several strategies have been designed to compare atomic coordination environments. In particular, the smooth overlap of atomic positions (SOAPs) has emerged as an elegant framework to obtain translation, rotation and permutation-invariant descriptors of groups of atoms, underlying the development of various classes of machine-learned inter-atomic potentials. Here we discuss how one can combine such local descriptors using a regularized entropy match (REMatch) approach to describe the similarity of both whole molecular and bulk periodic structures, introducing powerful metrics that enable the navigation of alchemical and structural complexities within a unified framework. Furthermore, using this kernel and a ridge regression method we can predict atomization energies for a database of small organic molecules with a mean absolute error below 1 kcal mol(-1), reaching an important milestone in the application of machine-learning techniques for the evaluation of molecular properties.