Domain-oriented Language Modeling with Adaptive Hybrid Masking and Optimal Transport Alignment

Domain-oriented Language Modeling with Adaptive Hybrid Masking and Optimal Transport Alignment
复制标题

DOI:
10.1145/3447548.3467215
复制
发表时间:
2021-08
期刊:
Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining
影响因子:
--
通讯作者:
Denghui Zhang;Zixuan Yuan;Yanchi Liu;Hao Liu;Fuzhen Zhuang;Hui Xiong;Haifeng Chen
Denghui Zhang;Zixuan Yuan;Yanchi Liu;Hao Liu;Fuzhen Zhuang;Hui Xiong;Haifeng Chen
中科院分区:
其他
文献类型:
--
作者:
Denghui Zhang;Zixuan Yuan;Yanchi Liu;Hao Liu;Fuzhen Zhuang;Hui Xiong;Haifeng Chen

文献摘要

相似文献

受 BERT 等预训练语言模型在广泛的自然语言处理 (NLP) 任务中取得成功的推动,最近的研究工作旨在使这些模型适应不同的应用领域。沿着这条线,现有的面向领域的模型主要遵循普通的 BERT 架构,并且可以直接使用领域语料库。然而,面向领域的任务通常需要对领域短语的准确理解,而现有的预训练方案很难捕获如此细粒度的短语级知识。此外,预训练模型的单词共现引导语义学习可以通过实体级关联知识在很大程度上得到增强。但与此同时,由于缺乏真实的字级对齐,存在引入噪声的风险。为了解决这些问题,我们提供了一种通用的面向领域的方法,利用辅助领域知识从两个方面改进现有的预训练框架。首先,为了有效地保存短语知识,我们构建了一个领域短语池作为辅助知识,同时引入自适应混合屏蔽模型来合并这些知识。它集成了单词学习和短语学习两种学习模式,并允许它们相互切换。其次,我们引入跨实体对齐,利用实体关联作为弱监督来增强预训练模型的语义学习。为了减轻这个过程中的潜在噪音,我们引入了一种基于可解释的最优传输的方法来指导对齐学习。对四个面向领域的任务的实验证明了我们框架的优越性。
Motivated by the success of pre-trained language models such as BERT in a broad range of natural language processing (NLP) tasks, recent research efforts have been made for adapting these models for different application domains. Along this line, existing domain-oriented models have primarily followed the vanilla BERT architecture and have a straightforward use of the domain corpus. However, domain-oriented tasks usually require accurate understanding of domain phrases, and such fine-grained phrase-level knowledge is hard to be captured by existing pre-training scheme. Also, the word co-occurrences guided semantic learning of pre-training models can be largely augmented by entity-level association knowledge. But meanwhile, there is a risk of introducing noise due to the lack of groundtruth word-level alignment. To address the issues, we provide a generalized domain-oriented approach, which leverages auxiliary domain knowledge to improve the existing pre-training framework from two aspects. First, to preserve phrase knowledge effectively, we build a domain phrase pool as auxiliary knowledge, meanwhile we introduce Adaptive Hybrid Masked Model to incorporate such knowledge. It integrates two learning modes, word learning and phrase learning, and allows them to switch between each other. Second, we introduce Cross Entity Alignment to leverage entity association as weak supervision to augment the semantic learning of pre-trained models. To alleviate the potential noise in this process, we introduce an interpretableOptimal Transport based approach to guide alignment learning. Experiments on four domain-oriented tasks demonstrate the superiority of our framework.