Anthropogenic biases in chemical reaction data hinder exploratory inorganic synthesis

Anthropogenic biases in chemical reaction data hinder exploratory inorganic synthesis
复制标题

DOI:
10.1038/s41586-019-1540-5
复制
发表时间:
2019-09-12
期刊:
影响因子:
64.8
通讯作者:
Schrier, Joshua
Schrier, Joshua
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Jia, Xiwen;Lynch, Allyson;Schrier, Joshua

文献摘要

被引文献

相似文献

大多数化学实验都是由人类科学家计划的,因此会受到各种人类认知偏见(1),化学(2)和社会影响(3)的影响。这些人为化学反应数据被广泛用于训练机器学习模型(4),这些模型用于预测有机(5)和无机(6,7)合成。然而,众所周知,社会偏见被编码在数据集中,并在机器学习模型中永久存在。在这里,我们确定尚未承认人为偏见的试剂选择和化学反应数据集的反应条件,使用数据挖掘和实验的组合。我们发现,胺模板金属氧化物的水热合成的晶体结构中的胺选择(9)遵循幂律分布,其中17%的胺反应物出现在79%的报告化合物中,与社会影响模型中的分布一致(10-12)。对未发表的历史实验室笔记本记录的分析表明,反应条件选择的分布也存在类似的偏倚。通过进行548个随机生成的实验,我们证明了反应物的受欢迎程度或反应条件的选择与反应的成功无关。我们表明,随机生成的实验更好地说明了与晶体形成兼容的参数选择范围。我们在较小的随机反应数据集上训练的机器学习模型优于在较大的人类选择的反应数据集上训练的模型,这表明了识别和解决科学数据中人为偏见的重要性。
Most chemical experiments are planned by human scientists and therefore are subject to a variety of human cognitive biases(1), heuristics(2) and social influences(3). These anthropogenic chemical reaction data are widely used to train machine-learning models(4) that are used to predict organic(5) and inorganic(6,7) syntheses. However, it is known that societal biases are encoded in datasets and are perpetuated in machine-learning models(8). Here we identify as-yet-unacknowledged anthropogenic biases in both the reagent choices and reaction conditions of chemical reaction datasets using a combination of data mining and experiments. We find that the amine choices in the reported crystal structures of hydrothermal synthesis of amine-templated metal oxides(9) follow a power-law distribution in which 17% of amine reactants occur in 79% of reported compounds, consistent with distributions in social influence models(10-12). An analysis of unpublished historical laboratory notebook records shows similarly biased distributions of reaction condition choices. By performing 548 randomly generated experiments, we demonstrate that the popularity of reactants or the choices of reaction conditions are uncorrelated to the success of the reaction. We show that randomly generated experiments better illustrate the range of parameter choices that are compatible with crystal formation. Machine-learning models that we train on a smaller randomized reaction dataset outperform models trained on larger human-selected reaction datasets, demonstrating the importance of identifying and addressing anthropogenic biases in scientific data.