Named Entity Recognition and Normalization Applied to Large-Scale Information Extraction from the Materials Science Literature

Named Entity Recognition and Normalization Applied to Large-Scale Information Extraction from the Materials Science Literature
复制标题

DOI:
10.1021/acs.jcim.9b00470
复制
发表时间:
2019-09-01
影响因子:
5.6
通讯作者:
Jain, A.
Jain, A.
中科院分区:
化学2区
文献类型:
--
作者:
Weston, L.;Tshitoyan, V;Jain, A.

文献摘要

被引文献

相似文献

在过去的几十年里,发表的材料科学文章的数量增加了许多倍。现在,材料发现管道中的一个主要瓶颈出现在将新结果与先前建立的文献联系起来。这个问题的一个潜在解决方案是将已发布文章的非结构化原始文本映射到允许编程查询的结构化数据库条目。为此,我们应用文本挖掘命名实体识别(NER)的大规模信息提取从出版的材料科学文献。NER模型经过训练,可以从材料科学文档中提取摘要级信息,包括无机材料提及、样品描述符、相标签、材料特性和应用,以及所使用的任何合成和表征方法。我们的分类器达到了87%的准确率(f(1)),并应用于327万材料科学摘要的信息提取。我们提取了超过8000万个与材料科学相关的命名实体,每个摘要的内容都以结构化格式表示为数据库条目。我们证明,简单的数据库查询可以用来回答复杂的“元问题”的已发表文献,以前需要费力,手动文献检索来回答。我们所有的数据和功能都可以在我们的Github(https://github.com/materialsintelligence/matscholar)和网站(http://www.example.com)上免费获得,我们希望这些结果能够加快未来材料科学发现的步伐。matscholar.com
The number of published materials science articles has increased manyfold over the past few decades. Now, a major bottleneck in the materials discovery pipeline arises in connecting new results with the previously established literature. A potential solution to this problem is to map the unstructured raw text of published articles onto structured database entries that allow for programmatic querying. To this end, we apply text mining with named entity recognition (NER) for large-scale information extraction from the published materials science literature. The NER model is trained to extract summary-level information from materials science documents, including inorganic material mentions, sample descriptors, phase labels, material properties and applications, as well as any synthesis and characterization methods used. Our classifier achieves an accuracy (f(1)) of 87%, and is applied to information extraction from 3.27 million materials science abstracts. We extract more than 80 million materials-science-related named entities, and the content of each abstract is represented as a database entry in a structured format. We demonstrate that simple database queries can be used to answer complex "meta-questions" of the published literature that would have previously required laborious, manual literature searches to answer. All of our data and functionality has been made freely available on our Github (https://github.com/materialsintelligence/matscholar) and website (http://matscholar.com), and we expect these results to accelerate the pace of future materials science discovery.