课题基金 / 基金详情

RI: Small: Improving Crowd-Sourced Annotation by Autonomous Intelligent Agents

RI: Small: Improving Crowd-Sourced Annotation by Autonomous Intelligent Agents
RI:小型:通过自主智能代理改进众包注释
批准号:
1420667
负责人:
Daniel Weld
金额:
$46.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-08-01 至 2018-07-31

项目摘要

项目成果

Daniel Weld的其他基金

相似基金

相关文献

中文摘要
翻译
可以说,有监督的机器学习方法是人工智能的最大成功案例,拥有深厚的理论基础和应用,从医疗诊断和科学数据分析,到电子商务推荐系统和信用卡欺诈检测。不幸的是,所有这些方法都需要标记的训练数据,而这些数据已经被人工注释-这是一个耗时且极其昂贵的过程。该项目将使用自动决策理论来控制注释过程,节省大量人力,并将机器学习的实际使用扩展到更广泛的社会问题。具体地说,这些方法解决了标签数据由大量人类注释员众包的情况,这些注释员的技能和错误率是可变的。该项目开发了新的控制算法,使学习者能够高效地要求特定的工作人员标记(或冗余地重新标记)特定的示例。为了测试他们方法的实用性,PI建立并进行了与Information Omnivore的研究,Information Omnivore是一个完全自主的代理,可以优化自然语言处理(NLP)训练数据的注释。通过不断向受薪工人和志愿公民科学家提出问题,杂食者将1)了解哪些问题难,哪些容易,2)了解各种工人的技能,3)决定问哪些工人的问题,以便在没有人的帮助的情况下最大限度地提高所学模型的准确性。除了对自动控制科学做出贡献外,杂食动物还将为两个重要的自然语言处理问题生成带标签的训练数据:命名实体链接(NEL)和信息提取(IE),极大地帮助了自然语言处理研究人员的社区。此外,研究人员还计划开展一系列外展活动,包括课程开发、参与太平洋科学中心的K12 Paws on Science项目,以及与华盛顿州学术红衫(STAR)工程项目的不同学生互动。PI提出的具体算法在几个方面值得注意。他们的决策理论优化框架实现了这样的直觉:(1)人们应该将更多或更好的员工分配到困难的问题上,(2)人们应该将精力从简单的问题或太难解决的任务中转移出来。自动化这一推理是困难的,因为问题难度和工人技能是潜在变量,因此代理必须面对探索/利用之间的权衡,因为它平衡了使其能够了解工人的能力的行动与产生高质量注释的最终目标。PI考虑两种情况:批注精确度的任务分配试图通过将工作人员批量分配到任务来最大化固定大小数据集的整体批注精确度。反动学习寻求通过平衡混合注释器请求来直接构造准确的ML分类器,以重新标记旧的或标记新的示例。在这两种情况下,他们都提出了一个基于决策理论方法的模型(例如,部分可观测马尔可夫决策过程(POMDP)和多武装强盗)。PI建议将他们的方法集成到Information Omnivore中,Information Omnivore是一个长期存在的软件代理,集成了计划和执行,在现实世界中行动,并学习其环境的模型。杂食动物将允许对他们的算法进行大规模的纬度研究,作为副产品,它将生成NLP训练数据,这将极大地帮助其他研究人员的大型社区。
英文摘要
Supervised machine learning methods are arguably the greatest success story for Artificial Intellitence with a deep underlying theory and applications ranging from medical diagnosis and scientific data analysis to ecommerce recommender systems and credit-card fraud detection. Unfortunately, all these methods require labeled training data, which has been annotated by a human --- a time consuming and extremely expensive process. This project will use automated decision theory to control the annotation process, saving significant amounts of human labor and extending the practical use of machine learning to a much broader array of societal problems. Specifically, the methods address the case where labeled data is crowd-sourced by a large number of human annotators whose skill and error rates are variable. The project develops new control algorithms that let the learner efficiently ask specific workers to label (or redundantly re-label) specific examples. To test the practicality of their methods, the PIs build and conduct studies with the Information Omnivore, a fully autonomous agent that optimizes the annotation of natural language processing (NLP) training data. By continuously posing questions to paid workers and volunteer citizen-scientists, the Omnivore 1) will learn which problems are hard and which are easy, 2) will learn about the skills of the various workers, 3) and will decide questions to ask which workers in order to maximize the accuracy of the learned model given scare human help. Besides contributing to the science of automated control, the Omnivore will generate labeled training data for two important NLP problems: named entity linking (NEL) and information extraction (IE), greatly helping the community of NLP researchers. Furthermore, the researchers plan a number of outreach efforts, including curriculum development, participation in the K12 Paws on Science program at the Pacific Science Center and interaction with the diverse students comprising the Washington STate Academic RedShirt (STARS) in Engineering program. The specific algorithms proposed by the PIs are notable in several respects. Their decision-theoretic optimization framework operationalizes intuitions like (1) one should assign more or better workers to hard problems and (2) one should redirect effort away from easy questions or from tasks that are too hard to solve. Automating this reasoning is hard because problem difficulty and worker skill are latent variables and thus the agent must confront an exploration / exploitation tradeoff as it balances actions that enable it to learn about the capabilities of workers with the ultimate goal of producing quality annotations. The PIs consider two cases: Task Allocation for Annotation Accuracy tries to maximize the overall annotation accuracy of a fixed size data set through batch assignment of workers to tasks. Re-Active Learning seeks instead to directly construct an accurate ML classifier through a balanced mix of annotator requests to re-label old or label new examples. In both cases they propose a model based on decision-theoretic methods (e.g., partially-observable Markov decision processes (POMDPs) and multi-armed bandits). The PIs propose to integrate their methods in the Information Omnivore, a long-lived software agent that integrates planning and execution, acts in the real world, and learns a model of its environment. The Omnivore will allow large-scale latitudinal studies of their algorithms, and as a byproduct will generate NLP training data that will greatly assist a large community of other researchers.
期刊论文(8)
专著(0)
科研奖励(0)
会议论文
DOI: --
发表时间: 2018
期刊: AAAI Conference on Human Computation
影响因子: --
作者: [C. Lin, Mausam]
通讯作者: C. Lin, Mausam
Intelligible Artificial Intelligence
可理解的人工智能
DOI: --
发表时间: 2018
期刊: March 2018
影响因子: --
作者: [D.S. Weld, G. Bansal]
通讯作者: D.S. Weld, G. Bansal
DOI: 10.18653/v1/n18-2058
发表时间: 2018-06
期刊: ArXiv
影响因子: --
作者: [James Ferguson;Colin Lockard;Daniel S. Weld;Hannaneh Hajishirzi]
通讯作者: James Ferguson;Colin Lockard;Daniel S. Weld;Hannaneh Hajishirzi
DOI: 10.1609/aaai.v32i1.11493
发表时间: 2018-04
期刊:
影响因子: --
作者: [Gagan Bansal;Daniel S. Weld]
通讯作者: Gagan Bansal;Daniel S. Weld
7
    CCRI: Research Infrastructure: NEW: Semantic Scholar Open Data Platform: Enabling Research Into Scientific Search and Discovery
    RAPID: Augmented Intelligence for Accelerating Covid-Related Scientific Discovery
    • 批准号:
      2040196
    • 项目类别:
      Standard Grant
    • 资助金额:
      $20.0万
    • 财政年份:
      2020
    • 负责人:
      Daniel Weld
    • 依托单位:
    RI: Small: Decision-Theoretic Control of Crowd-Sourced Workflows
    • 批准号:
      1016713
    • 项目类别:
      Standard Grant
    • 资助金额:
      $30.47万
    • 财政年份:
      2010
    • 负责人:
      Daniel Weld
    • 依托单位:
    RI: Small: Integrating Paradigms for Approximate Stochastic Planning
    • 批准号:
      1016465
    • 项目类别:
      Standard Grant
    • 资助金额:
      $45.05万
    • 财政年份:
      2010
    • 负责人:
      Daniel Weld
    • 依托单位:
    国内基金
    海外基金
    昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
    • 批准号:
    • 项目类别:
      省市级项目
    • 资助金额:
      --
    • 批准年份:
      2024
    • 负责人:
    • 依托单位:
    tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
    • 批准号:
    • 项目类别:
      省市级项目
    • 资助金额:
      10.0万元
    • 批准年份:
      2022
    • 负责人:
      张祥忠
    • 依托单位:
    Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
    Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
    • 批准号:
      31972324
    • 项目类别:
      面上项目
    • 资助金额:
      58.0万元
    • 批准年份:
      2019
    • 负责人:
      高学文
    • 依托单位: