BenchML: an extensible pipelining framework for benchmarking representations of materials and molecules at scale

BenchML: an extensible pipelining framework for benchmarking representations of materials and molecules at scale
复制标题

BenchML:一个可扩展的流水线框架,用于大规模地对材料和分子的表示进行基准测试

DOI:
--
复制
发表时间:
2021
期刊:
Machine Learning: Science and Technology
影响因子:
--
通讯作者:
Bingqing Cheng
Bingqing Cheng
中科院分区:
--
文献类型:
--
作者:
C. Poelking;Felix A Faber;Bingqing Cheng

文献摘要

被引文献

相似文献

我们引入了一个机器学习(ML)框架,用于对材料和分子数据集进行化学系统的不同表示的高通量基准测试。基准测试方法的指导原则是通过将模型复杂性限制在简单的回归方案中来评估原始描述符性能,同时实施最佳ML实践,允许无偏超参数优化,并通过学习曲线沿着沿着一系列同步训练测试分割来评估学习进度。由此产生的模型旨在作为基线,可以为未来的方法开发提供信息,此外还可以指示给定数据集的学习难度。通过对不同的物理化学、拓扑和几何表示的训练结果进行比较分析,我们深入了解了这些表示的相对优点以及它们之间的相互关系。
We introduce a machine-learning (ML) framework for high-throughput benchmarking of diverse representations of chemical systems against datasets of materials and molecules. The guiding principle underlying the benchmarking approach is to evaluate raw descriptor performance by limiting model complexity to simple regression schemes while enforcing best ML practices, allowing for unbiased hyperparameter optimization, and assessing learning progress through learning curves along series of synchronized train-test splits. The resulting models are intended as baselines that can inform future method development, in addition to indicating how easily a given dataset can be learnt. Through a comparative analysis of the training outcome across a diverse set of physicochemical, topological and geometric representations, we glean insight into the relative merits of these representations as well as their interrelatedness.