LEXPLAIN: Improving Model Explanations via Lexicon Supervision

LEXPLAIN: Improving Model Explanations via Lexicon Supervision
复制标题

DOI:
10.18653/v1/2023.starsem-1.19
复制
发表时间:
2023
期刊:
--
影响因子:
--
通讯作者:
Orevaoghene Ahia;Hila Gonen;Vidhisha Balachandran;Yulia Tsvetkov;Noah A. Smith
Orevaoghene Ahia;Hila Gonen;Vidhisha Balachandran;Yulia Tsvetkov;Noah A. Smith
中科院分区:
其他
文献类型:
--
作者:
Orevaoghene Ahia;Hila Gonen;Vidhisha Balachandran;Yulia Tsvetkov;Noah A. Smith

文献摘要

相似文献

对模型的预测阐明的模型正在成为NLP模型的额外输出,并在创建这些解释方面提出了挑战。通过明确监督它们的模型。我们的分析表明,在执行任务时,我们的方法还降低了伪造的相关性(即,关于非裔美国人的方言)。
Model explanations that shed light on the model’s predictions are becoming a desired additional output of NLP models, alongside their predictions. Challenges in creating these explanations include making them trustworthy and faithful to the model’s predictions. In this work, we propose a novel framework for guiding model explanations by supervising them explicitly. To this end, our method, LEXplain, uses task-related lexicons to directly supervise model explanations. This approach consistently improves the model’s explanations without sacrificing performance on the task, as we demonstrate on sentiment analysis and toxicity detection. Our analyses show that our method also demotes spurious correlations (i.e., with respect to African American English dialect) when performing the task, improving fairness.