CUE: An Uncertainty Interpretation Framework for Text Classifiers Built on Pre-Trained Language Models

CUE: An Uncertainty Interpretation Framework for Text Classifiers Built on Pre-Trained Language Models
复制标题

DOI:
10.48550/arxiv.2306.03598
复制
发表时间:
2023-06
期刊:
--
影响因子:
--
通讯作者:
Jiazheng Li;ZHAOYUE SUN;Bin Liang;Lin Gui;Yulan He
Jiazheng Li;ZHAOYUE SUN;Bin Liang;Lin Gui;Yulan He
中科院分区:
其他
文献类型:
--
作者:
Jiazheng Li;ZHAOYUE SUN;Bin Liang;Lin Gui;Yulan He

文献摘要

被引文献

相似文献

基于预训练的语言模型(PLM)建立的文本分类器在各种任务中取得了显着进步,包括情感分析,自然语言推断和提问。但是,这些分类器的不确定预测的发生在部署在实际应用中时的可靠性构成了挑战。为了了解PLM捕获的内容,已经大力努力设计各种探针。但是很少有研究研究影响基于PLM的分类器的预测不确定性的因素。在本文中,我们提出了一个名为CUE的新型框架,该框架旨在解释基于PLM模型的预测中固有的不确定性。特别是,我们首先通过差异自动编码器将PLM编码的表示形式映射到潜在空间。然后,我们通过扰动潜在空间来生成文本表示,从而导致预测不确定性波动。通过比较扰动和原始文本表示之间的预测不确定性差异,我们能够识别负责不确定性的潜在维度,然后随后追溯到有助于这种不确定性的输入特征。我们在四个基准数据集上进行了广泛的实验,其中包括语言可接受性分类,情感分类和自然语言推断,显示了我们提出的框架的可行性。我们的源代码可在以下网址提供:https://github.com/lijiazheng99/cue。
Text classifiers built on Pre-trained Language Models (PLMs) have achieved remarkable progress in various tasks including sentiment analysis, natural language inference, and question-answering. However, the occurrence of uncertain predictions by these classifiers poses a challenge to their reliability when deployed in practical applications. Much effort has been devoted to designing various probes in order to understand what PLMs capture. But few studies have delved into factors influencing PLM-based classifiers' predictive uncertainty. In this paper, we propose a novel framework, called CUE, which aims to interpret uncertainties inherent in the predictions of PLM-based models. In particular, we first map PLM-encoded representations to a latent space via a variational auto-encoder. We then generate text representations by perturbing the latent space which causes fluctuation in predictive uncertainty. By comparing the difference in predictive uncertainty between the perturbed and the original text representations, we are able to identify the latent dimensions responsible for uncertainty and subsequently trace back to the input features that contribute to such uncertainty. Our extensive experiments on four benchmark datasets encompassing linguistic acceptability classification, emotion classification, and natural language inference show the feasibility of our proposed framework. Our source code is available at: https://github.com/lijiazheng99/CUE.