Collective Bio-Entity Recognition in Scientific Documents using Hinge-Loss Markov Random Fields

Collective Bio-Entity Recognition in Scientific Documents using Hinge-Loss Markov Random Fields
复制标题

DOI:
--
复制
发表时间:
2020
期刊:
--
影响因子:
--
通讯作者:
A. Miller
A. Miller
中科院分区:
其他
文献类型:
--
作者:
A. Miller

文献摘要

相似文献

从科学文献中识别基因和蛋白质等生物实体对于回答问题和信息检索等下游任务至关重要。这项任务具有挑战性,因为相同的表面文本可以根据上下文指代基因或蛋白质。传统的方法,如Huang等人。[12]考虑周围文本中存在的单词来推断上下文。然而,他们没有考虑这些词的语义,这些词更好地表示为上下文词嵌入,如BERT [6]。另一方面,基于深度学习的方法未能利用科学文档的关系结构。我们引入了一种新的概率方法,该方法使用一类称为铰链损失马尔可夫随机场的无向图形模型对所有实体引用进行联合分类[1]。我们的方法可以将联合收割机的关系信息与基于嵌入的单词语义相结合。此外,我们的方法可以很容易地扩展到纳入新的信息来源。我们对JNLPBA共享任务语料库[4]的初步评估表明,我们的联合分类方法在F1得分上比传统机器学习方法和基于词嵌入的语义模型高出7.5%。
Identifying biological entities such as genes and proteins from scientific documents is crucial for further downstream tasks such as question answering and information retrieval. This task is challenging because the same surface text can refer either to a gene or a protein based on the context. Traditional approaches such as Huang et al. [12] consider the words present in the surrounding text to infer the context. However, they fail to consider the semantics of these words which are better represented by contextual word embeddings such as BERT [6]. Deep learning based approaches, on the other hand, fail to make use of the relational structure of scientific documents. We introduce a novel probabilistic approach that jointly classifies all entity references using a class of undirected graphical models called hinge-loss Markov random fields [1]. Our approach can combine relational information with embedding-based word semantics. Further, our approach can be easily extended to incorporate new sources of information. Our initial evaluation on the JNLPBA shared task corpus [4] shows that our joint classification approach outperforms both traditional machine learning approaches and semantic models based on word embeddings by up to 7.5% on F1 score.