Materials Synthesis Insights from Scientific Literature via Text Extraction and Machine Learning

Materials Synthesis Insights from Scientific Literature via Text Extraction and Machine Learning
复制标题

DOI:
10.1021/acs.chemmater.7b03500
复制
发表时间:
2017-11-14
影响因子:
8.6
通讯作者:
Olivetti, Elsa
Olivetti, Elsa
中科院分区:
材料科学2区
文献类型:
--
作者:
Kim, Edward;Huang, Kevin;Olivetti, Elsa

文献摘要

被引文献

相似文献

在过去的几年里,材料基因组计划(MGI)的努力已经产生了无数的例子,计算设计的材料在能源存储,催化,热电和储氢领域,以及用于筛选潜在的变革化合物的大型数据资源。因此,高通量材料设计的瓶颈已经转移到材料合成,这促使我们开发了一种方法,使用自然语言处理技术自动编译数万种学术出版物的材料合成参数。为了证明我们的框架的能力,我们检查了超过12000份手稿中各种金属氧化物的合成条件。然后,我们应用机器学习方法来预测通过水热方法合成二氧化钛纳米管所需的关键参数,并根据已知的机制验证这一结果。最后,我们通过使用机器学习模型来预测未包含在训练集中的材料系统的合成结果,从而超越启发式策略,展示了迁移学习的能力。
In the past several years, Materials Genome Initiative (MGI) efforts have produced myriad examples of computationally designed materials in the fields of energy storage, catalysis, thermoelectrics, and hydrogen storage as well as large data resources that are used to screen for potentially transformative compounds. The bottleneck in high-throughput materials design has thus shifted to materials synthesis, which motivates our development of a methodology to automatically compile materials synthesis parameters across tens of thousands of scholarly publications using natural language processing techniques. To demonstrate our framework's capabilities, we examine the synthesis conditions for various metal oxides across more than 12 thousand manuscripts. We then apply machine learning methods to predict the critical parameters needed to synthesize titania nanotubes via hydrothermal methods and verify this result against known mechanisms. Finally, we demonstrate the capacity for transfer learning by using machine learning models to predict synthesis outcomes on materials systems not included in the training set and thereby outperform heuristic strategies.