CI-P: Planning for Scalable Language Resource Creation through Novel Incentives and Crowdsourcing
CI-P: Planning for Scalable Language Resource Creation through Novel Incentives and Crowdsourcing
批准号:
1629923
负责人:
Christopher Cieri
金额:
$9.98万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-06-01 至 2018-04-30
中文摘要
例如,人类语言技术的进步使系统能够服从自然语言命令并作出实物响应,在许多语言对之间进行翻译,并总结多语言新闻。然而,这项技术的潜力在很大程度上仍未得到开发,因为推动发展的语言资源仍然远远不足。这个社区基础设施规划(CI-P)计划开始了构建基础设施的过程,通过采用在多个科学学科中证明有效的技术,不断开发高质量的语言资源。社交媒体、众包、有目的的游戏和公民科学向我们表明,在某些活动中,人力资源实际上是无限的。通过为人类贡献者提供适当的机会和激励,该项目大大提高了语言资源的开发,远远超出了直接资助所能产生的效果。通过消除对参与的限制,设计吸引多个社区的活动,该项目为公众创造了教育机会,包括学生和代表性不足的群体。数据规模和多样性的增加也有利于那些从事语言相关研究、教育和技术开发的人。不断增长的语言资源体的可用性将允许开发人员向世界上更大比例的人提供技术。这个项目是创建基础设施的第一步,能够通过无处不在、坚持不懈、全面注释、自动化培训和认证、适当的激励、任务工程和众包的变体来大量、持续地收集语言数据和判断。基于语言数据联盟的WebAnn框架,虚拟前端web服务器提供了多个界面来激励和设计来自目标群体的语言数据贡献:语言学家、公民科学家、游戏玩家和学生。收集和注释活动根据它们需要的技能被分析成组件任务,并适当地分配给使用不同工作流的不同工作人员。定制界面和新颖的激励策略相结合,使持续的、可扩展的数据收集和注释成为可能,从而为更广泛的计算机和信息科学与工程研究和教育社区提供了多种语言资源。
英文摘要
Advances in human language technologies enable systems that, for example, obey natural language commands and respond in kind, translate among many language pairs and summarize multilingual news. However, the technology's potential remains largely untapped because the linguistic resources that fuel development still fall far short of need. This community infrastructure planning (CI-P) initiative begins the process of building infrastructure to continuously develop high quality language resources, by employing techniques proven to work in multiple scientific disciplines. Social media, crowd-sourcing, games with a purpose and citizen science show us that human resources are effectively limitless for some activities. By offering human contributors appropriate opportunities and incentives, this project enhances language resource development well beyond what direct funding alone can produce. By removing constraints on participation, designing activities to appeal to multiple communities the project creates educational opportunities for the public including students and under-represented groups. The increase in scale and diversity of data also benefits those working in language related research, education and technology development. The availability of an ever-growing body of resources for an expanding range of languages will permit developers to supply technologies to a greater proportion of the world.This project is the first step in the creation of infrastructure capable of high volume, continuous collection of language data and judgments through: ubiquity, perseverance, comprehensive annotation, automated training and certification, appropriate incentives, task engineering and variants of crowdsourcing. Building upon Linguistic Data Consortium's WebAnn framework, virtual front end web servers provide multiple interfaces to incentivize and engineer linguistic data contributions from targeted groups: linguists, citizen scientists, game players and students. Collection and annotation activities are analyzed into component tasks according to the skills they require and are assigned as appropriate to different workforces using different workflows. The combination of customized interfaces and novel incentive strategies enables ongoing, scalable data collection and annotation resulting in diverse language resources available to the wider Computer and Information Science and Engineering research and education communities.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Workshop on Sociolinguistic Archive Preparation
-
批准号:1144480
-
项目类别:Standard Grant
-
资助金额:$2.45万
-
财政年份:2011
-
负责人:Christopher Cieri
-
依托单位:
CRI:CRD Collaborative Research: General Techniques for Creating Treebanks with Multiple Representations: A Large-Scale Russian Application
-
批准号:0708276
-
项目类别:Standard Grant
-
资助金额:$1.25万
-
财政年份:2007
-
负责人:Christopher Cieri
-
依托单位:
Networking Data Centers
-
批准号:9982201
-
项目类别:Standard Grant
-
资助金额:$15.45万
-
财政年份:2000
-
负责人:Christopher Cieri
-
依托单位:
海外基金