Few-Shot Fine-Grained Entity Typing with Automatic Label Interpretation and Instance Generation

Few-Shot Fine-Grained Entity Typing with Automatic Label Interpretation and Instance Generation
复制标题

DOI:
10.1145/3534678.3539443
复制
发表时间:
2022-06
期刊:
Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining
影响因子:
--
通讯作者:
Jiaxin Huang;Yu Meng;Jiawei Han
Jiaxin Huang;Yu Meng;Jiawei Han
中科院分区:
其他
文献类型:
--
作者:
Jiaxin Huang;Yu Meng;Jiawei Han

文献摘要

被引文献

相似文献

我们研究了少镜头细粒度实体类型(FET)问题,其中每个实体类型只给出几个带有上下文的注释实体提及。最近,通过将实体类型分类任务表述为“填空”问题,基于提示的调优在少数场景中表现出优于标准调优的性能。这允许有效地利用预训练语言模型(plm)强大的语言建模能力。尽管目前基于提示的调优方法取得了成功,但仍然存在两个主要挑战:(1)提示中的语言表达器要么是手工设计的,要么是从外部知识库构建的,没有考虑目标语料库和标签层次信息;(2)当前的方法主要利用PLMs的表示能力,但没有探索其通过广泛的通用领域预训练获得的生成能力。在这项工作中,我们提出了一个新的小样本FET框架,该框架由两个模块组成:(1)实体类型标签解释模块通过联合利用小样本实例和标签层次结构自动学习将类型标签与词汇表关联起来;(2)基于类型的上下文实例生成器根据给定的实例生成新的实例,以扩大训练集,从而更好地进行泛化。在三个基准数据集上,我们的模型明显优于现有方法。
We study the problem of few-shot Fine-grained Entity Typing (FET), where only a few annotated entity mentions with contexts are given for each entity type. Recently, prompt-based tuning has demonstrated superior performance to standard fine-tuning in few-shot scenarios by formulating the entity type classification task as a ''fill-in-the-blank'' problem. This allows effective utilization of the strong language modeling capability of Pre-trained Language Models (PLMs). Despite the success of current prompt-based tuning approaches, two major challenges remain: (1) the verbalizer in prompts is either manually designed or constructed from external knowledge bases, without considering the target corpus and label hierarchy information, and (2) current approaches mainly utilize the representation power of PLMs, but have not explored their generation power acquired through extensive general-domain pre-training. In this work, we propose a novel framework for few-shot FET consisting of two modules: (1) an entity type label interpretation module automatically learns to relate type labels to the vocabulary by jointly leveraging few-shot instances and the label hierarchy, and (2) a type-based contextualized instance generator produces new instances based on given instances to enlarge the training set for better generalization. On three benchmark datasets, our model outperforms existing methods by significant margins.