CAREER: Detecting, Understanding, and Fixing Vulnerabilities in Natural Language Processing Models
CAREER: Detecting, Understanding, and Fixing Vulnerabilities in Natural Language Processing Models
批准号:
2046873
负责人:
Sameer Singh
金额:
$50.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-07-01 至 2026-06-30
中文摘要
点击翻译按钮获取中文摘要
英文摘要
With recent advances in machine learning, models have achieved high accuracy on many challenging tasks in natural language processing (NLP) such as question answering, machine translation, and dialog agents, sometimes coming close to or beating human performance on these benchmarks. However, these NLP models often suffer from brittleness in many different ways: they latch onto erroneous artifacts, do not support natural variations in language, are not robust to adversarial attacks, and only work on a few domains. Existing pipelines for developing NLP models lack support for useful insights, and identifying bugs requires considerable effort from experts both in machine learning and the domain. This CAREER project develops several techniques to support this need for more robust training and evaluation pipelines for NLP, providing easy-to-use, scalable, and accurate mechanisms for identifying, understanding, and addressing NLP models' vulnerabilities. The developed methods will support diverse application areas such as conversational agents, sentiment classifiers, and abuse/hate speech detection. Further, the team engages with the developers of NLP models in academia and industry to develop a data science curriculum for K-12 education, particularly for students from underrepresented communities.Based on the notion of vulnerability as unexpected behavior on certain input transformations, the team will contribute across the following three thrusts. The first thrust identifies vulnerabilities by testing user-defined behaviors and searching over many possible vulnerabilities. In the second thrust, the investigators develop methods to understand the model's vulnerabilities by tracing the causes of errors to individual training data points and data artifacts. The last thrust will develop approaches to address vulnerabilities in models by directly injecting the vulnerability definitions into the model during training and using explanation-based annotations to supervise the models. These thrusts build upon the goals of behavioral testing, explanation-based interactions, and architecture agnosticism to support most current and future NLP models and applications.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(9)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Quantifying Social Biases Using Templates is Unreliable
使用模板量化社会偏见是不可靠的
DOI:
--
发表时间:
2022
期刊:
NeurIPS Workshop on Trustworthy and Socially Responsible Machine Learning (TSRML
影响因子:
--
作者:
[Seshadri, Preethi, Pezeshkpour, Pouya, Singh, Sameer]
通讯作者:
Singh, Sameer
Explaining machine learning models with interactive natural language conversations using TalkToModel
DOI:
10.1038/s42256-023-00692-8
发表时间:
2022-07
期刊:
Nature Machine Intelligence
影响因子:
23.8
作者:
[Dylan Slack;Satyapriya Krishna;Himabindu Lakkaraju;Sameer Singh]
通讯作者:
Dylan Slack;Satyapriya Krishna;Himabindu Lakkaraju;Sameer Singh
MISGENDERED: Limits of Large Language Models in Understanding Pronouns
性别错误:大型语言模型在理解代词方面的局限性
DOI:
10.18653/v1/2023.acl-long.293
发表时间:
2023
期刊:
Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers
影响因子:
--
作者:
[Hossain, Tamanna, Dev, Sunipa, Singh, Sameer]
通讯作者:
Singh, Sameer
DOI:
10.18653/v1/2022.findings-acl.153
发表时间:
2021-07
期刊:
ArXiv
影响因子:
--
作者:
[Pouya Pezeshkpour;Sarthak Jain;Sameer Singh;Byron C. Wallace]
通讯作者:
Pouya Pezeshkpour;Sarthak Jain;Sameer Singh;Byron C. Wallace
DOI:
--
发表时间:
2022
期刊:
影响因子:
--
作者:
[Sameer Singh]
通讯作者:
Sameer Singh
共 9 条
Collaborative Research: RI: Small: Post hoc Explanations in the Wild: Exposing Vulnerabilities and Ensuring Robustness
-
批准号:2008956
-
项目类别:Standard Grant
-
资助金额:$22.5万
-
财政年份:2020
-
负责人:Sameer Singh
-
依托单位:
CCRI: ENS: Machine Learning Democratization via a Linked, Annotated Repository of Datasets
-
批准号:1925741
-
项目类别:Standard Grant
-
资助金额:$179.3万
-
财政年份:2019
-
负责人:Sameer Singh
-
依托单位:
CRII: RI: Explaining Decisions of Black-box Models via Input Perturbations
-
批准号:1756023
-
项目类别:Standard Grant
-
资助金额:$17.49万
-
财政年份:2018
-
负责人:Sameer Singh
-
依托单位:
RI: Small: Modeling Multiple Modalities for Knowledge-Base Construction
-
批准号:1817183
-
项目类别:Standard Grant
-
资助金额:$44.8万
-
财政年份:2018
-
负责人:Sameer Singh
-
依托单位:
海外基金