课题基金 / 基金详情

A framework for evaluating and explaining the robustness of NLP models

A framework for evaluating and explaining the robustness of NLP models
评估和解释 NLP 模型稳健性的框架
批准号:
EP/X04162X/1
负责人:
Oana Cocarascu
金额:
$40.55万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2024
资助国家:
英国
项目状态:
未结题
起止时间:
2024 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
在NLP任务中评估受监督机器学习模型的泛化的标准实践是使用以前未见(即保持)的数据,并使用诸如准确性等各种度量来报告其性能。虽然基于保留数据的度量总结了模型的性能,但最终这些结果代表了基准的聚合统计数据,并不反映模型行为和稳健性在实际系统中应用的细微差别。我们提出了一个涉及参数和事实的NLP模型稳健性评估框架,其中包含对健壮性失败的解释,以支持系统和高效的评估。我们将开发新的方法来模拟来自现有数据集的真实世界文本,以帮助评估模型在野外部署时的稳定性和一致性。这些模拟方法将被用来通过在数据集以及捕捉语言模式的数据子集上进行基于文本的转换和分布平移来挑战NLP模型,以提供对真实世界语言现象的系统覆盖。此外,我们的框架将通过生成从各种数据集模拟和数据子集提取的词汇、形态和语法维度上的健壮性故障解释来深入了解模型的健壮性,从而与当前仅提供度量来量化健壮性的方法不同。我们将集中在两个NLP研究领域,论点挖掘和事实验证,然而,几种模拟方法和健壮性解释也可以扩展到其他NLP任务。
英文摘要
The standard practice for evaluating the generalisation of supervised machine learning models in NLP tasks is to use previously unseen (i.e. held-out) data and report the performance on it using various metrics such as accuracy. Whilst metrics reported on held-out data summarise a model's performance, ultimately these results represent aggregate statistics on benchmarks and do not reflect the nuances in model behaviour and robustness when applied in real-world systems.We propose a robustness evaluation framework for NLP models concerned with arguments and facts, which encompasses explanations for robustness failures to support systematic and efficient evaluation. We will develop novel methods for simulating real-world texts stemming from existing datasets, to help evaluate the stability and consistency of models when deployed in the wild. The simulation methods will be used to challenge NLP models through text-based transformations and distribution shifts on datasets as well as on data sub-sets that capture linguistic patterns, to provide a systematic coverage of real-world linguistic phenomena. Furthermore, our framework will shed insights into a model's robustness by generating explanations for robustness failures along the lexical, morphological, and syntactic dimensions, extracted from the various dataset simulations and data sub-sets, thus departing from current approaches that solely provide a metric to quantify robustness. We will focus on two NLP research areas, argument mining and fact verification, however, several simulation methods and the robustness explanations are also scalable to other NLP tasks.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金