Self-learning entropic population annealing for interpretable materials design

Self-learning entropic population annealing for interpretable materials design
复制标题

用于可解释材料设计的自学习熵布居退火

DOI:
10.1039/d1dd00043h
复制
发表时间:
2022
期刊:
Digital Discovery
影响因子:
--
通讯作者:
Tsuda Koji
Tsuda Koji
中科院分区:
--
文献类型:
--
作者:
Li Jiawen;Zhang Jinzhe;Tamura Ryo;Tsuda Koji

文献摘要

相似文献

在自动材料设计中,从黑盒优化中获得的样本为科学家提供了一个获得新知识的有吸引力的机会。通常进行样品的统计分析,例如,来发现关键描述符由于大多数黑盒优化算法是有偏的采样器,事后分析可能会导致误导性的结论。为了科普这个问题,我们提出了一种新的方法,称为自学习熵种群退火(SLEPA),结合熵采样和代理机器学习模型。SLEPA的样本带有权重,可正确估计目标属性和感兴趣描述符的联合分布。在短肽设计中,将SLEPA与纯黑盒优化在估计目标性质的多个阈值处的残基分布方面进行了比较。虽然黑盒优化在目标属性的尾部更好,但SLEPA在广泛的阈值范围内更好。我们的研究结果表明,如何调和统计一致性与有效的优化材料发现。
In automatic materials design, samples obtained from black-box optimization offer an attractive opportunity for scientists to gain new knowledge. Statistical analyses of the samples are often conducted, e.g., to discover key descriptors. Since most black-box optimization algorithms are biased samplers, post hoc analyses may result in misleading conclusions. To cope with the problem, we propose a new method called self-learning entropic population annealing (SLEPA) that combines entropic sampling and a surrogate machine learning model. Samples of SLEPA come with weights to estimate the joint distribution of the target property and a descriptor of interest correctly. In short peptide design, SLEPA was compared with pure black-box optimization in estimating the residue distributions at multiple thresholds of the target property. While black-box optimization was better at the tail of the target property, SLEPA was better for a wide range of thresholds. Our result shows how to reconcile statistical consistency with efficient optimization in materials discovery.