Can Language Models be Biomedical Knowledge Bases?

Can Language Models be Biomedical Knowledge Bases?
复制标题

语言模型可以成为生物医学知识库吗?

DOI:
--
复制
发表时间:
2021
期刊:
Conference on Empirical Methods in Natural Language Processing
影响因子:
--
通讯作者:
Jaewoo Kang
Jaewoo Kang
中科院分区:
--
文献类型:
--
作者:
Mujeen Sung;Jinhyuk Lee;Sean S. Yi;Minji Jeon;Sungdong Kim;Jaewoo Kang

文献摘要

参考文献

被引文献

相似文献

预先训练的语言模型(LM)在解决各种自然语言处理(NLP)任务中变得无处不在。人们越来越感兴趣的是这些LM包含什么知识,以及我们如何提取这些知识,将LM视为知识库(KB)。虽然在一般领域中已经有很多关于探测LM的工作,但是很少有人关注这些强大的LM是否可以用作特定于领域的知识库。为此,我们创建了BioLAMA基准,它由49K生物医学事实知识三元组组成,用于探测生物医学LM。我们发现,生物医学LM最近提出的探测方法可以实现高达18.51% Acc@5检索生物医学知识。虽然这似乎是有希望的任务难度,我们的详细分析表明,大多数预测是高度相关的提示模板没有任何主题,从而产生类似的结果,每个关系,并阻碍其能力被用作特定领域的知识库。我们希望BioLAMA可以作为一个具有挑战性的基准生物医学事实探测。
Pre-trained language models (LMs) have become ubiquitous in solving various natural language processing (NLP) tasks. There has been increasing interest in what knowledge these LMs contain and how we can extract that knowledge, treating LMs as knowledge bases (KBs). While there has been much work on probing LMs in the general domain, there has been little attention to whether these powerful LMs can be used as domain-specific KBs. To this end, we create the BioLAMA benchmark, which is comprised of 49K biomedical factual knowledge triples for probing biomedical LMs. We find that biomedical LMs with recently proposed probing methods can achieve up to 18.51% Acc@5 on retrieving biomedical knowledge. Although this seems promising given the task difficulty, our detailed analyses reveal that most predictions are highly correlated with prompt templates without any subjects, hence producing similar results on each relation and hindering their capabilities to be used as domain-specific KBs. We hope that BioLAMA can serve as a challenging benchmark for biomedical factual probing.
DOI: 10.1162/tacl_a_00324
发表时间: 2020-01-01
影响因子: 10.9
作者:
Jiang, Zhengbao;Xu, Frank F.;Neubig, Graham
通讯作者: Neubig, Graham