Exploring Pre-Trained Language Models to Build Knowledge Graph for Metal-Organic Frameworks (MOFs)

Exploring Pre-Trained Language Models to Build Knowledge Graph for Metal-Organic Frameworks (MOFs)
复制标题

DOI:
10.1109/bigdata55660.2022.10020568
复制
发表时间:
2022-12
期刊:
2022 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Yuan An;Jane Greenberg;Xiaohua Hu;Alexander Kalinowski;Xiao Fang;Xintong Zhao;Scott McClellan;F. Uribe-Romo;Kyle Langlois;Jacob Furst;Diego A. Gómez-Gualdrón;Fernando Fajardo-Rojas;Katherine Ardila;S. Saikin;Corey A. Harper;Ron Daniel
Yuan An;Jane Greenberg;Xiaohua Hu;Alexander Kalinowski;Xiao Fang;Xintong Zhao;Scott McClellan;F. Uribe-Romo;Kyle Langlois;Jacob Furst;Diego A. Gómez-Gualdrón;Fernando Fajardo-Rojas;Katherine Ardila;S. Saikin;Corey A. Harper;Ron Daniel
中科院分区:
其他
文献类型:
--
作者:
Yuan An;Jane Greenberg;Xiaohua Hu;Alexander Kalinowski;Xiao Fang;Xintong Zhao;Scott McClellan;F. Uribe-Romo;Kyle Langlois;Jacob Furst;Diego A. Gómez-Gualdrón;Fernando Fajardo-Rojas;Katherine Ardila;S. Saikin;Corey A. Harper;Ron Daniel

文献摘要

相似文献

构建知识图是一个耗时且成本高昂的过程,它通常应用复杂的自然语言处理(NLP)方法从文本语料库中提取知识图三元组。预训练的大型语言模型(PLM)已经成为一种重要的方法,为一系列AI应用提供了现成的知识。然而,目前还不清楚从PLM构建特定领域的知识图是否可行。受知识图加速数据驱动材料发现的能力的激励,我们探索了一组最先进的预训练通用和特定领域的语言模型,以提取金属有机框架(MOF)的知识三元组。我们为1248个已发布的MOF同义词创建了一个具有7个关系的知识图基准。我们的实验结果表明,特定领域的PLM始终优于通用PLM预测MOF相关的三元组。然而,总体基准测试结果表明,使用目前的PLM来创建特定领域的知识图仍然远远不够实用,因此需要为材料科学中的特定应用开发更有能力和知识的预训练语言模型。
Building a knowledge graph is a time-consuming and costly process which often applies complex natural language processing (NLP) methods for extracting knowledge graph triples from text corpora. Pre-trained large Language Models (PLM) have emerged as a crucial type of approach that provides readily available knowledge for a range of AI applications. However, it is unclear whether it is feasible to construct domain-specific knowledge graphs from PLMs. Motivated by the capacity of knowledge graphs to accelerate data-driven materials discovery, we explored a set of state-of-the-art pre-trained general-purpose and domain-specific language models to extract knowledge triples for metal-organic frameworks (MOFs). We created a knowledge graph benchmark with 7 relations for 1248 published MOF synonyms. Our experimental results showed that domain-specific PLMs consistently outperformed the general-purpose PLMs for predicting MOF related triples. The overall benchmarking results, however, show that using the present PLMs to create domain-specific knowledge graphs is still far from being practical, motivating the need to develop more capable and knowledgeable pre-trained language models for particular applications in materials science.