课题基金 / 基金详情

CAREER: High-Agreement Crowdsourcing for Difficult Language-Understanding Tasks

CAREER: High-Agreement Crowdsourcing for Difficult Language-Understanding Tasks
职业:针对困难的语言理解任务的高度一致的众包
批准号:
2046556
负责人:
Samuel Bowman
金额:
$55.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-10-01 至 2026-09-30

项目摘要

项目成果

Samuel Bowman的其他基金

相似基金

相关文献

中文摘要
翻译
当工程师为问题回答等语言问题构建现代人工智能(AI)系统时,他们使用实例数据集来教系统如何解决问题,而不是直接对系统进行编程。这些例子的数据集通常是通过群体工作来收集的,在这种工作中,大量非专家被雇用来提出问题的范例答案、文献的范例摘要等。让不同的人提供付费数据是为了让快速建立专门的语言技术系统成为可能,并确保它们可以涵盖广泛的语言风格,但这在实践中并不总是奏效:群体工作的设置往往迫使参与者快速而草率地工作,产生的数据在教机器做我们想做的事情时效率低下。该奖项支持旨在解决这一问题的研究,方法是开发和评估群工培训、反馈和奖金支付的最佳实践,以帮助群工数据集创建者发展专业技能并产生更好的数据,从而产生真正有效的语言技术。该项目奖还将支持同时培养新科学家和工程师的努力,包括针对高级技术生的编程和针对该领域新手的外联活动。从技术上讲,该项目将为阅读理解问题回答、共指关系解决和自然语言推理等自然语言理解任务建立一套有科学依据的众包数据收集做法,重点放在能够确保结果数据多样化、具有挑战性和高质量的方法上,面对主观性和合法注释者分歧带来的障碍。主要的实验是分离几种新的数据收集技术的效果,包括在众包数据收集中使用的培训、反馈和激励结构。一个互补的线索将评估和改进任务设计,目标是确定最能分离和加强模型理解文本和与文本推理的能力的任务公式,了解现有任务的大型实验调查。随之而来的教育项目将通过研讨会和讲授的研究方法课程,扩大研究导师的过程,以接触到纽约大学多样化且合格的本科生和研究生群体的更大比例。随附的外展计划将支持为对人工智能和语言技术职业暂时感兴趣的早期本科生开发一个周期性的研讨会系列,特别是从计算机领域代表性较低的群体招聘。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
When engineers build modern artificial intelligence (AI) systems for language problems like question answering, they use datasets of examples to teach the systems how to solve the problem, rather than programming the systems directly. These datasets of examples are often collected through crowd work, where a large population of non-specialists are hired to come up with example answers to questions, example summaries of documents, or the like. Having a diverse group of people provide data for pay is meant to make it possible to build specialized language technology systems quickly, and to ensure that they can cover a wide range of styles of language, but this has not always worked well in practice: Crowd work is often set up in a way that forces participants to work quickly and sloppily, and produces data that’s ineffective at teaching machines to do what we want. This award supports research that aims to fix this, by developing and evaluating best practices for crowd worker training, feedback, and bonus pay to help crowd-worker dataset creators develop professional skills and produce better data that will lead to truly effective language technologies. The project award will also support parallel efforts at training new scientists and engineers, including programming targeting advanced technical students and outreach events targeting newcomers to the field.Technically, the project will establish a scientifically-grounded set of practices for crowdsourced data collection for natural language understanding tasks like reading comprehension question answering, coreference resolution, and natural language inference, with a focus on methods that can ensure that the resulting data is diverse, challenging, and high-quality in the face of obstacles posed by subjectivity and legitimate annotator disagreements. The main experiments to isolate the effect of several novel techniques for data collection, covering the training, feedback, and incentive structures used in crowdsourced data collection. A complementary thread will evaluate and refine task designs with the goal of identifying the task formulations that best isolate and reinforce model abilities to understand and reason with texts, informed by large experimental surveys of existing tasks. The accompanying education program will scale up processes for research mentorship to reach a larger fraction of the diverse and qualified undergraduate and graduate student population at New York University, both through seminars and taught research methods courses. The accompanying outreach plan will support the development of a recurring workshop series for early-year undergraduates tentatively interested in careers in AI and language technology, recruiting especially from groups underrepresented in computing.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(11)
专著(0)
科研奖励(0)
会议论文
DOI: 10.18653/v1/2022.acl-long.516
发表时间: 2021-10
期刊:
影响因子: --
作者: [Sam Bowman]
通讯作者: Sam Bowman
DOI: 10.18653/v1/2021.blackboxnlp-1.42
发表时间: 2021-09
期刊:
影响因子: --
作者: [Jason Phang;Haokun Liu;Samuel R. Bowman]
通讯作者: Jason Phang;Haokun Liu;Samuel R. Bowman
Single-Turn Debate Does Not Help Humans Answer Hard Reading-Comprehension Questions
单轮辩论无助于人类回答困难的阅读理解问题
DOI: 10.18653/v1/2022.lnls-1.3
发表时间: 2022
期刊: Proceedings of the First Workshop on Learning with Natural Language Supervision
影响因子: --
作者: [Parrish, Alicia, Trivedi, Harsh, Perez, Ethan, Chen, Angelica, Nangia, Nikita, Phang, Jason, Bowman, Samuel]
通讯作者: Bowman, Samuel
DOI: 10.18653/v1/2022.findings-acl.165
发表时间: 2021-10
期刊:
影响因子: --
作者: [Alicia Parrish;Angelica Chen;Nikita Nangia;Vishakh Padmakumar;Jason Phang;Jana Thompson;Phu Mon Htut;Sam Bowman]
通讯作者: Alicia Parrish;Angelica Chen;Nikita Nangia;Vishakh Padmakumar;Jason Phang;Jana Thompson;Phu Mon Htut;Sam Bowman
10
    CRII: RI: Can Low-Bias Machine Learners Acquire English Grammar? Deep Learning and Linguistic Acceptability
    • 批准号:
      1850208
    • 项目类别:
      Standard Grant
    • 资助金额:
      $17.49万
    • 财政年份:
      2019
    • 负责人:
      Samuel Bowman
    • 依托单位:
    The 2018 NAACL Student Research Workshop
    • 批准号:
      1803423
    • 项目类别:
      Standard Grant
    • 资助金额:
      $1.5万
    • 财政年份:
      2018
    • 负责人:
      Samuel Bowman
    • 依托单位:
    海外基金