课题基金 / 基金详情

EAGER: CISE/IIS/RI/Program Element 7495: Crowdsourcing for NLP: Exploring Two Approaches

EAGER: CISE/IIS/RI/Program Element 7495: Crowdsourcing for NLP: Exploring Two Approaches
EAGER:CISE/IIS/RI/Program Element 7495:NLP 众包:探索两种方法
批准号:
0947841
负责人:
Collin Baker
金额:
$20.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2009
资助国家:
美国
项目状态:
已结题
起止时间:
2009-08-15 至 2013-01-31
关键词:

项目摘要

项目成果

Collin Baker的其他基金

相似基金

相关文献

中文摘要
翻译
“众包”指的是利用“众智”,即结合非专家的大量判断,为复杂问题提供可靠的答案。在自然语言处理(NLP)领域,对句子进行标注以显示它们表达了什么事件(以及句子的哪些部分表达了哪些参与者)是一项非常复杂的任务。例如,句子“Maria搭乘公交车从家里到她的办公室”应该被认定为Ride_Vehicle事件,“Maria”是mover,“the bus”是车辆,“from home”是来源,“to her office”是目的;自然语言处理系统还应该能够识别相同的事件和句子中的相同参与者,但大多数当前的系统不能。FrameNet(http://framenet.icsi.berkeley.edu)正在建立一个包含成百上千种事件类型(称为“语义框架”)的词汇数据库以及每种事件类型的注释句子的实例,该数据库可用于训练自然语言处理系统。但专家对句子的注释既慢又贵;这个项目正在测试众包能否加快这类数据库的创建,特别是通过探索两种众包技术,看看哪种更适合这些任务:(1)在线游戏,玩家们竞争,看谁能快速准确地注释(类似于“冗长”游戏)和(2)一个系统,在这个系统中,使用亚马逊的“机械土耳其人”(www.mturk.com),人们可以拿到少量的钱来完成这些任务。如果成功,这些技术可以被用来为新的NLP系统建立更好的数据库,这些系统真正理解“谁对谁做了什么”,从而改进了问题回答和网络搜索。
英文摘要
"Crowdsourcing" is the idea of using the "wisdom of crowds", that is, combining large numbers of judgments by non-experts, to produce reliable answers to complex problems. In the field of natural language processing(NLP), annotating sentences to show what events they express (and which parts of the sentence express which participants) is such a complex task. For example, the sentence "Maria rides the bus from home to her office" should be recognized as a Ride_vehicle event, with "Maria" as Mover, "the bus" as the Vehicle, "from home" as the Source and "to her office" as the Goal; NLP systems should also be able to recognize the same event with the same participants in the sentence "Maria's bus ride from home to her office takes 40 minutes", but most current systems cannot.FrameNet (http://framenet.icsi.berkeley.edu) is building a lexical database of hundreds of event types (called "semantic frames") and examples of each in annotated sentences, which can be used to train NLP systems. But expert annotation of sentences is slow and expensive; this project is testing whether crowdsourcing can speed up the creation of such databases, specifically by exploring two crowdsourcing techniques to see which works better for these tasks: (1) online games, where players compete to see who can annotate rapidly and accurately (similar to the "Verbosity" game) and (2) a system in which people are paid small amounts of money to complete such tasks, using Amazon's "Mechanical Turk" (www.mturk.com). If successful, these techniques could be used to build better databases for new NLP systems that really understand "who did what to whom", thus improving question answering and web searching.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Berkeley FrameNet Website Migration
CI-NEW: Multilingual FrameNet: A Resource Enabling Cross-Lingual Research for the Natural Language Processing Community
CI-P: Planning for a Multilingual FrameNet Lexical Resource
FrameNet Workshop: Developing New NLP Applications
海外基金