Domain-Invariant Feature Progressive Distillation with Adversarial Adaptive Augmentation for Low-Resource Cross-Domain NER

Domain-Invariant Feature Progressive Distillation with Adversarial Adaptive Augmentation for Low-Resource Cross-Domain NER
复制标题

DOI:
10.1145/3570502
复制
发表时间:
2022-12
影响因子:
2
通讯作者:
Tao Zhang;Congying Xia;Zhiwei Liu;Shu Zhao;Hao Peng;Philip S. Yu
Tao Zhang;Congying Xia;Zhiwei Liu;Shu Zhao;Hao Peng;Philip S. Yu
中科院分区:
计算机科学4区
文献类型:
--
作者:
Tao Zhang;Congying Xia;Zhiwei Liu;Shu Zhao;Hao Peng;Philip S. Yu

文献摘要

相似文献

考虑到命名实体识别(NER)中标注代价大的问题,跨域命名实体识别通过传递高资源领域的知识,使命名实体识别能够在低资源目标领域中实现标注数据少或没有标注数据的情况。然而,不同域之间的差异导致域转移问题,并阻碍了低资源场景下的跨域NER的性能。在这篇文章中,我们首先提出了一种对抗性自适应增强,将对抗性策略集成到多任务学习器中,以增强和限定领域自适应数据。我们提取自适应数据的域不变特征,以弥合跨域的差距,同时减轻标签稀疏性问题。因此,本文中的另一个重要组件是渐进域不变特征提取框架。该框架中的多粒度MMD(最大均值离散)方法可以提取多层次的领域不变特征,并通过对抗性自适应数据实现跨领域的知识转移。高级知识蒸馏(KD)模式通过强大的预训练语言模型和多层次领域不变特征逐步进行领域适应。在四个英语和两个中文基准上进行的大量对比实验表明,对抗增强和从高资源域到低资源目标域的有效适应的重要性。与两个香草和四个最新的基线的比较表明,最先进的性能和优越性面临的零资源和最小资源的情况。
Considering the expensive annotation in Named Entity Recognition (NER), Cross-domain NER enables NER in low-resource target domains with few or without labeled data, by transferring the knowledge of high-resource domains. However, the discrepancy between different domains causes the domain shift problem and hampers the performance of cross-domain NER in low-resource scenarios. In this article, we first propose an adversarial adaptive augmentation, where we integrate the adversarial strategy into a multi-task learner to augment and qualify domain adaptive data. We extract domain-invariant features of the adaptive data to bridge the cross-domain gap and alleviate the label-sparsity problem simultaneously. Therefore, another important component in this article is the progressive domain-invariant feature distillation framework. A multi-grained MMD (Maximum Mean Discrepancy) approach in the framework to extract the multi-level domain invariant features and enable knowledge transfer across domains through the adversarial adaptive data. Advanced Knowledge Distillation (KD) schema processes progressively domain adaptation through the powerful pre-trained language models and multi-level domain invariant features. Extensive comparative experiments over four English and two Chinese benchmarks show the importance of adversarial augmentation and effective adaptation from high-resource domains to low-resource target domains. Comparison with two vanilla and four latest baselines indicates the state-of-the-art performance and superiority confronted with both zero-resource and minimal-resource scenarios.