课题基金 / 基金详情

Fine-grained Interpretation of Deep Neural Network Models of NLP

Fine-grained Interpretation of Deep Neural Network Models of NLP
NLP深度神经网络模型细粒度解读
批准号:
RGPIN-2022-03943
负责人:
Sajjad, Hassan
金额:
$2.11万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31

项目摘要

项目成果

Sajjad, Hassan的其他基金

相似基金

相关文献

中文摘要
翻译
我们解释深度神经网络(DNN)的能力是迈向可解释人工智能的关键一步,以确保模型的公平性,并实现学习更好泛化的模型。在这项建议中,我们致力于通过提供对DNN模型表示的细粒度解释来提高DNN模型的透明度。细粒度解释的目标是某些知识在网络中的编码方式。它回答了这样的问题:神经元在网络中扮演什么角色?在神经元中学习了什么任务和语言知识?了解神经元的角色及其在网络中的重要性,为建立透明、公平、可解释和有效的模型铺平了道路。在这个提案中,我们将对DNN的细粒度解释做出多方面的贡献。更具体地说,我们将提出用于解释的数据集和算法。我们将致力于解释模型预测的因果解释。此外,我们将开发通过神经元分析实现的应用程序。具体项目总结如下:概念数据集:当前DNN解释工作探索人类定义的语言概念,如名词、动词等是否在模型中学习。然而,他们不分析模型了解到的关于任务和语言的其他潜在概念。该项目旨在准备一个以模型为中心的概念数据集。我们将使用高维空间中的一组词的上下文表示来识别它们,并将在某些标准下手动对它们进行注释。得到的数据集是一个新颖的多方面数据,并将支持以模型为中心的解释。解释算法:神经元本质上是多变量的,模型由于训练和架构选择而表现出冗余。如何识别冗余地学习一个概念的神经元?如何找到一组共同工作的神经元来代表一个概念?神经元分析的常见方法大多忽略了这些问题,或者需要仔细评估才能证明它们的有效性。该项目旨在改进神经元分析的方法。我们将探索各种多变量特征选择方法,如(稀疏)组套索来分析神经元。因果解释:对于可解释的人工智能来说,概念对于一个类或特定预测的重要性是必不可少的。这个项目的目标是对预测进行基于概念的因果解释。我们将使用基于梯度和扰动的归因算法来识别预测的显著神经元,并将其与神经元分析联系起来,以生成基于概念的预测解释。应用:在这个项目中,我们将探索神经元分析的应用,更具体地说,控制模型的行为和领域适应。其想法是根据感兴趣的概念(如性别或领域)识别神经元,并开发在测试时操纵这些神经元的方法,以便以理想的方式控制模型的行为。
英文摘要
Our ability to interpret Deep neural networks (DNNs) is an essential step towards explainable AI, to ensure fairness of models, and in achieving models that learn better generalization. In this proposal, we work towards increasing the transparency of DNN models by providing fine-grained interpretation of their representation. Fine-grained interpretation targets how certain knowledge is encoded in the network. It answers questions such as, what is the role of neurons in the network? What task and language knowledge is learned in neurons?, Knowing the role of neurons and their importance in the network paves the way towards transparent, fair, explainable and efficient models. In this proposal, we will make multiple contributions to fine-grained interpretation of DNNs. More specifically, we will propose datasets and algorithms for interpretation. We will work towards causal interpretation that explains the prediction of a model. Moreover, we will develop applications that are enabled by neuron analysis. The specific projects are summarized below: Concept Dataset: Current work in DNN interpretation probes whether human-defined linguistic concepts such as noun, verb, etc. are learned in a model. However, they do not analyze what other latent concepts a model has learned about the task and the language. This project aims at preparing a model-centric concept dataset. We will identify group of words in high dimensional space using their contextualized representation and will manually annotate them under certain criteria. The resulting dataset is a novel multi-facet data and will enable model-centered interpretation. Interpretation Algorithms: Neurons are multivariate in nature and models exhibit redundancy due to training and architectural choices. How to identify neurons that redundantly learn a concept? How to find a group of neurons that work together to represent a concept? Common approaches to neuron analysis largely ignore these questions or require careful evaluation to prove their efficacy. This project aims at improving methods of neuron analysis. We will explore various multivariate feature selection methods such as (sparse) group lasso to analyze neurons. Causal Interpretation: The importance of a concept for a class or for a specific prediction is essential for explainable AI. This project targets concept-based causal interpretation of prediction. We will use gradient and perturbation-based attribution algorithms to identify salient neurons for a prediction and will connect it with neuron analysis to generate concept-based explanations of predictions. Applications: In this project, we will explore applications of neuron analysis, more specifically controlling model's behavior and domain adaptation. The idea is to identify neurons with respect to concepts of interest e.g. gender or domain and develop methods to manipulate those neurons at test time in order to control the behavior of the model in a desirable way.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Fine-grained Interpretation of Deep Neural Network Models of NLP
  • 批准号:
    DGECR-2022-00384
  • 项目类别:
    Discovery Launch Supplement
  • 资助金额:
    $0.91万
  • 财政年份:
    2022
  • 负责人:
    Sajjad, Hassan
  • 依托单位:
海外基金