课题基金 / 基金详情

CAREER: Learning and Using Community-Driven Natural Language Processing Models

CAREER: Learning and Using Community-Driven Natural Language Processing Models
职业:学习和使用社区驱动的自然语言处理模型
批准号:
2145357
负责人:
Anthony Rios
金额:
$55.16万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-06-01 至 2027-05-31

项目摘要

项目成果

Anthony Rios的其他基金

相似基金

相关文献

中文摘要
翻译
该奖项的全部或部分资金来自《2021年美国救援计划法案》(公法117-2)。人们对将自然语言处理(NLP)应用于广泛的任务越来越感兴趣,包括但不限于健康、在线审核和教育。与NLP相关的研究一般都集中在统一适用于每个人的大型模式,而不是他们的写作风格和社会规范,因此,假设一个一刀切的解决方案。然而,基于NLP的模式并不适用于所有社区,因为不同的写作风格(例如方言)和主题讨论的选择(例如体育与科技)。此外,不同社区的社会规范各不相同,这使得一些NLP模型最初的预期用途可能无关紧要。因此,如果不直接考虑社区,将相同的NLP模型应用于每个人可能会造成伤害。因此,研究人员和实践者必须在部署NLP模型之前对社区数据进行评估。他们还必须与社区合作,根据社区的社会规范和需求确定技术是否健全。该项目将解决两个关键问题:“利益相关者如何知道该模型在投入生产时是否会损害特定的社区?”以及“哪些特定于社区的语言模式会导致各种NLP模型中的错误?”通过回答这些问题,这个项目打算开发工具来帮助社区参与技术开发过程,这将使他们能够决定一项特定的技术是否与社区相关。最后,该项目还将通过培训当地高中教师学习社区驱动的自然语言规划,为圣安东尼奥的高中生创建基于标准的课程,这将潜在地帮助当地学生更好地考虑自然语言规划在他们生活中的应用。总的来说,该项目将为研究人员和自然语言编程从业者提供更好的理解,使他们能够更好地应用和开发适用于小型社区的自然语言规划模式,而不是专注于一刀切的框架。具体地说,为了实现这一目标,该项目有三个目标。目标1将确定策略,以发现特定于社区的语言与NLP模型性能之间的相关性。这一目标将导致更好地理解社区特定应用的NLP模型的归纳偏差。目标2将创建一个工具,通过帮助社区特定的利益攸关方确定应用特定的NLP模式可能导致的潜在的好的和有害的结果,从而促进参与性的NLP设计。更重要的是,这一目标将帮助决策者决定何时应该或不应该在他们的社区部署NLP。目标3将确定针对特定社区改进NLP模式的方法。目标是确定将社区的未标记数据合并到现有已标记NLP数据集中的方法,以提高社区特定模型的性能。最后,该项目将通过发布实现该奖项产生的工具和技术的开源软件来影响更广泛的NLP社区。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This award is funded in whole or in part under the American Rescue Plan Act of 2021 (Public Law 117-2).There is a growing interest in applying Natural Language Processing (NLP) to a wide array of tasks, including, but not limited to, health, online moderation, and education. NLP-related research has generally focused on large models uniformly applied to everyone independent of their writing style and social norms, thus, assuming a one-size-fits-all solution. Nevertheless, NLP-based models do not perform equally for all communities because of different writing styles (e.g., dialects) and choices of topical discussion (e.g., Sports vs. Technology). Furthermore, social norms vary between communities, making the original intended use of some NLP models potentially irrelevant. Hence, applying the same NLP model to everyone may cause harm if communities are not directly considered. Therefore, researchers and practitioners must evaluate NLP models on community data before they deploy them. They must also work with communities to determine whether the technology is sound given the community's social norms and needs. This project will address two critical questions: "How can stakeholders know whether the model will harm specific communities when put into production?" and "What community-specific language patterns cause errors in various NLP models?". By answering these questions, this project intends to develop tools to help communities participate in the technology development process, which will enable them to decide whether a specific technology is relevant to the community or not. Finally, this project will also create standards-based lessons for high school students in San Antonio by training local high-school teachers in community-driven NLP, which will potentially help local students better consider NLP applications in their lives.Overall, this project will provide researchers and NLP practitioners with a better understanding of applying and developing NLP models for small communities rather than focusing on a one-size-fits-all framework. Specifically, to address this goal, this project has three objectives. Objective 1 will identify strategies to find correlations between community-specific language and NLP model performance. This objective will result in a better understanding of the inductive biases of NLP models for community-specific applications. Objective 2 will create a tool that can facilitate participatory NLP design by helping community-specific stakeholders identify potential good and harmful outcomes that may be caused by applying a specific NLP model. More importantly, the objective will help decision-makers decide when NLP should or should not be deployed in their communities. Objective 3 will identify methods to improve NLP models for specific communities. The goal is to identify methods to incorporate a community's unlabeled data into existing labeled NLP datasets to improve community-specific model performance. Finally, the project will impact the broader NLP community via the release of open-source software that implements the tools and techniques this award generates.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(9)
专著(0)
科研奖励(0)
会议论文
DOI: 10.48550/arxiv.2212.12799
发表时间: 2022-12
期刊: ArXiv
影响因子: --
作者: [Xingmeng Zhao;A. Niazi;Anthony Rios]
通讯作者: Xingmeng Zhao;A. Niazi;Anthony Rios
A marker-based neural network system for extracting social determinants of health
用于提取健康社会决定因素的基于标记的神经网络系统
DOI: 10.1093/jamia/ocad041
发表时间: 2023
期刊: Journal of the American Medical Informatics Association
影响因子: 6.4
作者: [Zhao, Xingmeng, Rios, Anthony]
通讯作者: Rios, Anthony
Measuring Geographic Performance Disparities of Offensive Language Classifiers
衡量攻击性语言分类器的地理表现差异
DOI: --
发表时间: 2022
期刊: COLING
影响因子: --
作者: [Brandon, Lwowski, Rad, Paul, Rios, Anthony]
通讯作者: Rios, Anthony
Linguistic Elements of Engaging Customer Service Discourse on Social Media
在社交媒体上参与客户服务对话的语言元素
DOI: --
发表时间: 2022
期刊: Proceedings of the Fifth Workshop on Natural Language Processing and Computational Social Science (NLP+CSS
影响因子: --
作者: [Singh, Sonam, Rios, Anthony]
通讯作者: Rios, Anthony
8
    CRII: SCH: A Computational Framework for Fair Public Health-Related Decisions
    • 批准号:
      1947697
    • 项目类别:
      Standard Grant
    • 资助金额:
      $17.48万
    • 财政年份:
      2020
    • 负责人:
      Anthony Rios
    • 依托单位:
    国内基金
    海外基金
    Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
    Understanding structural evolution of galaxies with machine learning
    • 批准号:
    • 项目类别:
      省市级项目
    • 资助金额:
      10.0万元
    • 批准年份:
      2022
    • 负责人:
      Nicola Rosario Napolitano
    • 依托单位:
    煤矿安全人机混合群智感知任务的约束动态多目标Q-learning进化分配
    • 批准号:
      --
    • 项目类别:
      青年科学基金项目
    • 资助金额:
      30万元
    • 批准年份:
      2022
    • 负责人:
      吉建娇
    • 依托单位:
    基于领弹失效考量的智能弹药编队短时在线Q-learning协同控制机理
    • 批准号:
      62003314
    • 项目类别:
      青年科学基金项目
    • 资助金额:
      24.0万元
    • 批准年份:
      2020
    • 负责人:
      沈剑
    • 依托单位: