Collective Bio-Entity Recognition in Scientific Documents using Hinge-Loss Markov Random Fields
Collective Bio-Entity Recognition in Scientific Documents using Hinge-Loss Markov Random Fields
复制标题
DOI:
--
复制
发表时间:
2020
期刊:
影响因子:
--
通讯作者:
A. Miller
中科院分区:
文献类型:
--
作者:
A. Miller
Identifying biological entities such as genes and proteins from scientific documents is crucial for further downstream tasks such as question answering and information retrieval. This task is challenging because the same surface text can refer either to a gene or a protein based on the context. Traditional approaches such as Huang et al. [12] consider the words present in the surrounding text to infer the context. However, they fail to consider the semantics of these words which are better represented by contextual word embeddings such as BERT [6]. Deep learning based approaches, on the other hand, fail to make use of the relational structure of scientific documents. We introduce a novel probabilistic approach that jointly classifies all entity references using a class of undirected graphical models called hinge-loss Markov random fields [1]. Our approach can combine relational information with embedding-based word semantics. Further, our approach can be easily extended to incorporate new sources of information. Our initial evaluation on the JNLPBA shared task corpus [4] shows that our joint classification approach outperforms both traditional machine learning approaches and semantic models based on word embeddings by up to 7.5% on F1 score.