Few-Shot Named Entity Recognition via Meta-Learning

Few-Shot Named Entity Recognition via Meta-Learning
复制标题

DOI:
10.1109/tkde.2020.3038670
复制
发表时间:
2022-09-01
影响因子:
8.9
通讯作者:
Wang, Hao
Wang, Hao
中科院分区:
计算机科学2区
文献类型:
--
作者:
Li, Jing;Chiu, Billy;Wang, Hao

文献摘要

被引文献

相似文献

N-way K-shot设置下的Few-shot学习(即N个类别中每个类别有K个带注释的样本)在关系提取(例如FewRel)和图像分类(例如Mini-ImageNet)中得到了广泛的研究。命名实体识别(NER)通常被定义为一个序列标记问题,其中实体类别固有地纠缠在一起,因为句子中的实体编号和类别是事先未知的,这使得N-way K-shot NER问题到目前为止尚未被探索。在本文中,我们首先正式定义了一个更合适的N-way K-shot设置。然后,我们提出了一种新的元学习方法FEWNER。FEWNER将整个网络分为任务独立部分和任务特定部分。在FEWNER的训练过程中,任务无关部分是跨多个任务进行元学习的,而任务特定部分是在低维空间中针对每个单独的任务进行元学习的。在测试时,FEWNER保持与任务无关的部分固定,并通过梯度下降只更新特定于任务的部分来适应新的任务,从而降低了过拟合的可能性,提高了计算效率。与以隐式方式(即依赖大规模语料库)获得可迁移性的预训练语言模型(如BERT和ELMo)相比,FEWNER通过元学习明确地优化了“学习快速适应”的能力。结果表明,在域内交叉类型、跨域内类型和跨域交叉类型3个适应实验中,FEwNER在9种基线方法的基础上达到了最先进的性能。
Few-shot learning under the N-way K-shot setting (i.e., K annotated samples for each of N classes) has been widely studied in relation extraction (e.g., FewRel) and image classification (e.g., Mini-ImageNet). Named entity recognition (NER) is typically framed as a sequence labeling problem where the entity classes are inherently entangled together because the entity number and classes in a sentence are not known in advance, leaving the N-way K-shot NER problem so far unexplored. In this paper, we first formally define a more suitable N-way K-shot setting for NER. Then we propose FEWNER, a novel meta-learning approach for few-shot NER. FEWNER separates the entire network into a task-independent part and a task-specific part. During training in FEWNER, the task-independent part is meta-learned across multiple tasks and the task-specific part is learned for each individual task in a low-dimensional space. At test time, FEWNER keeps the task-independent part fixed and adapts to a new task via gradient descent by updating only the task-specific part, resulting in it being less prone to overfitting and more computationally efficient. Compared with pre-trained language models (e.g., BERT and ELMo) which obtain the transferability in an implicit manner (i.e., relying on large-scale corpora), FEWNER explicitly optimizes the capability of "learning to adapt quickly" through meta-learning. The results demonstrate that FEwNER achieves state-of-the-art performance against nine baseline methods by significant margins on three adaptation experiments (i.e., intra-domain cross-type, cross-domain intra-type and cross-domain cross-type).