CAREER: Learning and Using Community-Driven Natural Language Processing Models
CAREER: Learning and Using Community-Driven Natural Language Processing Models
批准号:
2145357
负责人:
Anthony Rios
金额:
$55.16万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-06-01 至 2027-05-31
中文摘要
该奖项全部或部分由《2021年美国救援计划法案》(公法117-2)资助。人们对将自然语言处理(NLP)应用于广泛的任务越来越感兴趣,包括但不限于健康、在线调节和教育。与nlp相关的研究通常集中在大型模型上,这些模型统一适用于每个人,而不受他们的写作风格和社会规范的影响,因此,假设一个通用的解决方案。然而,由于不同的写作风格(例如,方言)和主题讨论的选择(例如,体育与技术),基于nlp的模型并不适用于所有社区。此外,社会规范因社区而异,使得一些NLP模型的最初预期用途可能不相关。因此,如果不直接考虑社区,将相同的NLP模型应用于每个人可能会造成伤害。因此,研究人员和实践者必须在部署NLP模型之前对社区数据进行评估。他们还必须与社区合作,根据社区的社会规范和需求来确定技术是否合理。该项目将解决两个关键问题:“涉众如何知道模型投入生产时是否会损害特定社区?”以及“哪些特定社区的语言模式会导致各种NLP模型中的错误?”通过回答这些问题,该项目打算开发工具来帮助社区参与技术开发过程,这将使他们能够决定特定技术是否与社区相关。最后,该项目还将通过培训当地高中教师社区驱动的NLP,为圣安东尼奥的高中生创建基于标准的课程,这将有可能帮助当地学生更好地考虑在他们的生活中应用NLP。总的来说,这个项目将为研究人员和NLP从业者提供更好的理解应用和开发小型社区的NLP模型,而不是专注于一个通用的框架。具体来说,为了实现这个目标,这个项目有三个目标。目标1将确定寻找社区特定语言和NLP模型性能之间相关性的策略。这一目标将有助于更好地理解NLP模型在社区特定应用中的归纳偏差。目标2将创建一个工具,通过帮助特定社区的利益相关者识别应用特定NLP模型可能导致的潜在好的和有害的结果,从而促进参与式NLP设计。更重要的是,这个目标将帮助决策者决定什么时候应该或不应该在他们的社区中部署NLP。目标3将确定针对特定社区改进NLP模型的方法。目标是确定将社区未标记数据纳入现有标记NLP数据集的方法,以提高社区特定模型的性能。最后,该项目将通过发布实现该奖项产生的工具和技术的开源软件来影响更广泛的NLP社区。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This award is funded in whole or in part under the American Rescue Plan Act of 2021 (Public Law 117-2).There is a growing interest in applying Natural Language Processing (NLP) to a wide array of tasks, including, but not limited to, health, online moderation, and education. NLP-related research has generally focused on large models uniformly applied to everyone independent of their writing style and social norms, thus, assuming a one-size-fits-all solution. Nevertheless, NLP-based models do not perform equally for all communities because of different writing styles (e.g., dialects) and choices of topical discussion (e.g., Sports vs. Technology). Furthermore, social norms vary between communities, making the original intended use of some NLP models potentially irrelevant. Hence, applying the same NLP model to everyone may cause harm if communities are not directly considered. Therefore, researchers and practitioners must evaluate NLP models on community data before they deploy them. They must also work with communities to determine whether the technology is sound given the community's social norms and needs. This project will address two critical questions: "How can stakeholders know whether the model will harm specific communities when put into production?" and "What community-specific language patterns cause errors in various NLP models?". By answering these questions, this project intends to develop tools to help communities participate in the technology development process, which will enable them to decide whether a specific technology is relevant to the community or not. Finally, this project will also create standards-based lessons for high school students in San Antonio by training local high-school teachers in community-driven NLP, which will potentially help local students better consider NLP applications in their lives.Overall, this project will provide researchers and NLP practitioners with a better understanding of applying and developing NLP models for small communities rather than focusing on a one-size-fits-all framework. Specifically, to address this goal, this project has three objectives. Objective 1 will identify strategies to find correlations between community-specific language and NLP model performance. This objective will result in a better understanding of the inductive biases of NLP models for community-specific applications. Objective 2 will create a tool that can facilitate participatory NLP design by helping community-specific stakeholders identify potential good and harmful outcomes that may be caused by applying a specific NLP model. More importantly, the objective will help decision-makers decide when NLP should or should not be deployed in their communities. Objective 3 will identify methods to improve NLP models for specific communities. The goal is to identify methods to incorporate a community's unlabeled data into existing labeled NLP datasets to improve community-specific model performance. Finally, the project will impact the broader NLP community via the release of open-source software that implements the tools and techniques this award generates.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(9)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.48550/arxiv.2212.12799
发表时间:
2022-12
期刊:
ArXiv
影响因子:
--
作者:
[Xingmeng Zhao;A. Niazi;Anthony Rios]
通讯作者:
Xingmeng Zhao;A. Niazi;Anthony Rios
A marker-based neural network system for extracting social determinants of health
用于提取健康社会决定因素的基于标记的神经网络系统
DOI:
10.1093/jamia/ocad041
发表时间:
2023
期刊:
Journal of the American Medical Informatics Association
影响因子:
6.4
作者:
[Zhao, Xingmeng, Rios, Anthony]
通讯作者:
Rios, Anthony
DOI:
--
发表时间:
2022
期刊:
COLING
影响因子:
--
作者:
[Brandon, Lwowski, Rad, Paul, Rios, Anthony]
通讯作者:
Rios, Anthony
Linguistic Elements of Engaging Customer Service Discourse on Social Media
在社交媒体上参与客户服务对话的语言元素
DOI:
--
发表时间:
2022
期刊:
Proceedings of the Fifth Workshop on Natural Language Processing and Computational Social Science (NLP+CSS
影响因子:
--
作者:
[Singh, Sonam, Rios, Anthony]
通讯作者:
Rios, Anthony
DOI:
10.48550/arxiv.2403.17363
发表时间:
2024-03
期刊:
ArXiv
影响因子:
--
作者:
[Nima Ebadi;Kellen Morgan;Adrian Tan;Billy Linares;Sheri Osborn;Emma Majors;Jeremy Davis;Anthony Rios]
通讯作者:
Nima Ebadi;Kellen Morgan;Adrian Tan;Billy Linares;Sheri Osborn;Emma Majors;Jeremy Davis;Anthony Rios
共 8 条
CRII: SCH: A Computational Framework for Fair Public Health-Related Decisions
-
批准号:1947697
-
项目类别:Standard Grant
-
资助金额:$17.48万
-
财政年份:2020
-
负责人:Anthony Rios
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
-
批准号:--
-
项目类别:合作创新研究团队
-
资助金额:--
-
批准年份:2024
-
负责人:姚韬
-
依托单位:
Understanding structural evolution of galaxies with machine learning
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:Nicola Rosario Napolitano
-
依托单位:
煤矿安全人机混合群智感知任务的约束动态多目标Q-learning进化分配
-
批准号:--
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2022
-
负责人:吉建娇
-
依托单位:
基于领弹失效考量的智能弹药编队短时在线Q-learning协同控制机理
-
批准号:62003314
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:沈剑
-
依托单位:
集成上下文张量分解的e-learning资源推荐方法研究
-
批准号:61902016
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2019
-
负责人:万珊珊
-
依托单位:
具有时序迁移能力的Spiking-Transfer learning (脉冲-迁移学习)方法研究
-
批准号:61806040
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2018
-
负责人:解修蕊
-
依托单位:
基于Deep-learning的三江源区冰川监测动态识别技术研究
-
批准号:51769027
-
项目类别:地区科学基金项目
-
资助金额:38.0万元
-
批准年份:2017
-
负责人:张大奇
-
依托单位:
具有时序处理能力的Spiking-Deep Learning(脉冲深度学习)方法研究
-
批准号:61573081
-
项目类别:面上项目
-
资助金额:64.0万元
-
批准年份:2015
-
负责人:屈鸿
-
依托单位:
基于有向超图的大型个性化e-learning学习过程模型的自动生成与优化
-
批准号:61572533
-
项目类别:面上项目
-
资助金额:66.0万元
-
批准年份:2015
-
负责人:孙雪冬
-
依托单位:
E-Learning中学习者情感补偿方法的研究
-
批准号:61402392
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2014
-
负责人:秦继伟
-
依托单位: