Leveraging Symbolic Knowledge Bases for Commonsense Natural Language Inference Using Pattern Theory

Leveraging Symbolic Knowledge Bases for Commonsense Natural Language Inference Using Pattern Theory
复制标题

DOI:
10.1109/tpami.2023.3287837
复制
发表时间:
2023-06
影响因子:
23.6
通讯作者:
Sathyanarayanan N. Aakur;Sudeep Sarkar
Sathyanarayanan N. Aakur;Sudeep Sarkar
中科院分区:
计算机科学1区
文献类型:
--
作者:
Sathyanarayanan N. Aakur;Sudeep Sarkar

文献摘要

相似文献

常识自然语言推理(CNLI)任务的目的是选择最有可能的后续声明,以普通的,日常的事件和事实的上下文描述。当前跨任务转移CNLI模型学习的方法需要来自新任务的许多标记数据。本文提出了一种方法,以减少这种需要额外的注释训练数据的新任务,利用符号知识库,如概念网。我们制定了一个教师-学生框架的混合符号神经推理,与大规模的符号知识库作为教师和训练有素的CNLI模型作为学生。这种混合蒸馏过程包括两个步骤。第一步是符号推理过程。给定一组未标记的数据,我们使用基于Grenander模式理论的溯因推理框架来创建弱标记数据。模式理论是一种基于能量的图形概率框架,用于在具有不同依赖结构的随机变量之间进行推理。在第二步中,弱标记数据,沿着的标记数据的一小部分,被用来转移学习CNLI模型到新的任务。目标是减少所需的标记数据的比例。我们通过使用三个公开可用的数据集(OpenBookQA,SWAG和HellaSWAG)并评估代表不同任务的三个CNLI模型(BERT,LSTM和ESIM)来证明我们方法的有效性。我们表明,平均而言,我们在没有标记数据的情况下实现了完全监督BERT模型最高性能的63%。只需1,000个标记样本,我们就可以将此性能提高到72%。有趣的是,在没有训练的情况下,教师机制本身就有很大的推理能力。模式理论框架在OpenBookQA上达到了32.7%的准确率,远远超过了基于transformer的模型,如GPT(26.6%),GPT-2(30.2%)和BERT(27.1%)。我们证明,该框架可以推广到成功地训练神经CNLI模型,在无监督和半监督学习设置下使用知识蒸馏。我们的研究结果表明,它优于所有无监督和弱监督基线以及一些早期的监督方法,同时提供具有竞争力的性能与完全监督基线。此外,我们表明,溯因学习框架可以适用于其他下游任务,如无监督语义文本相似性,无监督情感分类和零镜头文本分类,而无需对框架进行重大修改。最后,用户研究表明,生成的解释,提高其可解释性提供了关键的见解,其推理机制。
The commonsense natural language inference (CNLI) tasks aim to select the most likely follow-up statement to a contextual description of ordinary, everyday events and facts. Current approaches to transfer learning of CNLI models across tasks require many labeled data from the new task. This paper presents a way to reduce this need for additional annotated training data from the new task by leveraging symbolic knowledge bases, such as ConceptNet. We formulate a teacher-student framework for mixed symbolic-neural reasoning, with the large-scale symbolic knowledge base serving as the teacher and a trained CNLI model as the student. This hybrid distillation process involves two steps. The first step is a symbolic reasoning process. Given a collection of unlabeled data, we use an abductive reasoning framework based on Grenander's pattern theory to create weakly labeled data. Pattern theory is an energy-based graphical probabilistic framework for reasoning among random variables with varying dependency structures. In the second step, the weakly labeled data, along with a fraction of the labeled data, is used to transfer-learn the CNLI model into the new task. The goal is to reduce the fraction of labeled data required. We demonstrate the efficacy of our approach by using three publicly available datasets (OpenBookQA, SWAG, and HellaSWAG) and evaluating three CNLI models (BERT, LSTM, and ESIM) that represent different tasks. We show that, on average, we achieve 63% of the top performance of a fully supervised BERT model with no labeled data. With only 1,000 labeled samples, we can improve this performance to 72%. Interestingly, without training, the teacher mechanism itself has significant inference power. The pattern theory framework achieves 32.7% accuracy on OpenBookQA, outperforming transformer-based models such as GPT (26.6%), GPT-2 (30.2%), and BERT (27.1%) by a significant margin. We demonstrate that the framework can be generalized to successfully train neural CNLI models using knowledge distillation under unsupervised and semi-supervised learning settings. Our results show that it outperforms all unsupervised and weakly supervised baselines and some early supervised approaches, while offering competitive performance with fully supervised baselines. Additionally, we show that the abductive learning framework can be adapted for other downstream tasks, such as unsupervised semantic textual similarity, unsupervised sentiment classification, and zero-shot text classification, without significant modification to the framework. Finally, user studies show that the generated interpretations enhance its explainability by providing key insights into its reasoning mechanism.