Semi-Supervised Bidirectional Long Short-Term Memory and Conditional Random Fields Model for Named-Entity Recognition Using Embeddings from Language Models Representations.

Semi-Supervised Bidirectional Long Short-Term Memory and Conditional Random Fields Model for Named-Entity Recognition Using Embeddings from Language Models Representations.
复制标题

使用语言模型表示的嵌入进行命名实体识别的半监督双向长短期记忆和条件随机场模型

DOI:
10.3390/e22020252
复制
发表时间:
2020-02-22
期刊:
Entropy (Basel, Switzerland)
影响因子:
--
通讯作者:
Chen J
Chen J
中科院分区:
其他
文献类型:
--
作者:
Zhang M;Geng G;Chen J

文献摘要

参考文献

相似文献

越来越受欢迎的在线博物馆显著改变了人们获取文化知识的方式。这些在线博物馆一直在产生大量的文物数据。近年来,研究人员使用可以自动提取复杂特征并具有丰富表示能力的深度学习模型来实现命名实体识别(NER)。然而,文物领域缺乏标记数据,使得依赖于标记数据的深度学习模型难以获得优异的性能。为了解决这个问题,本文提出了一种半监督深度学习模型SCRNER(Semi-supervised model for Cultural Religious ' Named Entity Recognition),该模型利用双向长短期记忆(BiLSTM)和条件随机场(CRF)模型,通过很少标记的数据和大量未标记的数据进行训练,以获得有效的性能。为了满足半监督的样本选择,我们提出了一种重复标记(relabeled)策略来选择高置信度的样本,以迭代地扩大训练集。此外,本文还采用基于语言模型(埃尔莫)表示的嵌入方法,动态获取单词表示作为模型的输入,以解决文物领域中文物边界模糊和文本具有中国特色的问题。实验结果表明,我们提出的模型,训练有限的标记数据,实现了有效的性能在文物命名实体识别任务。
Increasingly, popular online museums have significantly changed the way people acquire cultural knowledge. These online museums have been generating abundant amounts of cultural relics data. In recent years, researchers have used deep learning models that can automatically extract complex features and have rich representation capabilities to implement named-entity recognition (NER). However, the lack of labeled data in the field of cultural relics makes it difficult for deep learning models that rely on labeled data to achieve excellent performance. To address this problem, this paper proposes a semi-supervised deep learning model named SCRNER (Semi-supervised model for Cultural Relics’ Named Entity Recognition) that utilizes the bidirectional long short-term memory (BiLSTM) and conditional random fields (CRF) model trained by seldom labeled data and abundant unlabeled data to attain an effective performance. To satisfy the semi-supervised sample selection, we propose a repeat-labeled (relabeled) strategy to select samples of high confidence to enlarge the training set iteratively. In addition, we use embeddings from language model (ELMo) representations to dynamically acquire word representations as the input of the model to solve the problem of the blurred boundaries of cultural objects and Chinese characteristics of texts in the field of cultural relics. Experimental results demonstrate that our proposed model, trained on limited labeled data, achieves an effective performance in the task of named entity recognition of cultural relics.
使用条件随机字段的 Hadoop 识别生物医学命名实体
DOI: 10.1109/tpds.2014.2368568
发表时间: 2015-11-01
影响因子: 5.3
作者:
Li, Kenli;Ai, Wei;Hwang, Kai
通讯作者: Hwang, Kai
DOI: 10.1177/0735633117752614
发表时间: 2019-04-01
影响因子: 4.8
作者:
Livieris, Ioannis E.;Drakopoulou, Konstantina;Pintelas, Panagiotis
通讯作者: Pintelas, Panagiotis
DOI: 10.1007/978-3-642-24797-2
发表时间: 1997-11-15
期刊: NEURAL COMPUTATION
影响因子: 2.9
作者:
Hochreiter, S;Schmidhuber, J
通讯作者: Schmidhuber, J
DOI: 10.31449/inf.v43i2.2217
发表时间: 2019-06-01
影响因子: --
作者:
Livieris, Ioannis
通讯作者: Livieris, Ioannis