CI-P: Planning for Scalable Language Resource Creation through Novel Incentives and Crowdsourcing
CI-P: Planning for Scalable Language Resource Creation through Novel Incentives and Crowdsourcing
批准号:
1629923
负责人:
Christopher Cieri
金额:
$9.98万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-06-01 至 2018-04-30
中文摘要
例如,人类语言技术的进步使系统能够服从自然语言命令并以实物回应,在许多语言对之间进行翻译并总结多语言新闻。然而,这项技术的潜力在很大程度上仍未得到开发,因为推动发展的语言资源仍然远远不能满足需要。这个社区基础设施规划(CI-P)计划开始了基础设施建设的过程,通过采用在多个科学学科中被证明有效的技术,不断开发高质量的语言资源。社交媒体、众包、有目的的游戏和公民科学向我们表明,人力资源在某些活动中实际上是无限的。通过为人类贡献者提供适当的机会和激励措施,该项目促进了语言资源的开发,远远超出了直接资助的范围。通过消除对参与的限制,设计吸引多个社区的活动,该项目为包括学生和代表性不足的群体在内的公众创造了教育机会。数据规模和多样性的增加也使从事语言相关研究、教育和技术开发的人员受益。不断增长的语言资源将使开发人员能够向世界上更大的地区提供技术。该项目是创建能够通过以下方式连续收集大量语言数据和判断的基础设施的第一步:无处不在、坚持不懈、全面注释、自动化培训和认证、适当激励、任务工程和众包变体。基于语言数据联盟的WebAnn框架,虚拟前端Web服务器提供多个接口,以激励和设计目标群体(语言学家,公民科学家,游戏玩家和学生)的语言数据贡献。收集和注释活动根据其所需的技能被分析为组件任务,并根据需要分配给使用不同工作流的不同工作人员。定制界面和新颖的激励策略相结合,使持续的,可扩展的数据收集和注释,从而为更广泛的计算机和信息科学与工程研究和教育社区提供多样化的语言资源。
英文摘要
Advances in human language technologies enable systems that, for example, obey natural language commands and respond in kind, translate among many language pairs and summarize multilingual news. However, the technology's potential remains largely untapped because the linguistic resources that fuel development still fall far short of need. This community infrastructure planning (CI-P) initiative begins the process of building infrastructure to continuously develop high quality language resources, by employing techniques proven to work in multiple scientific disciplines. Social media, crowd-sourcing, games with a purpose and citizen science show us that human resources are effectively limitless for some activities. By offering human contributors appropriate opportunities and incentives, this project enhances language resource development well beyond what direct funding alone can produce. By removing constraints on participation, designing activities to appeal to multiple communities the project creates educational opportunities for the public including students and under-represented groups. The increase in scale and diversity of data also benefits those working in language related research, education and technology development. The availability of an ever-growing body of resources for an expanding range of languages will permit developers to supply technologies to a greater proportion of the world.This project is the first step in the creation of infrastructure capable of high volume, continuous collection of language data and judgments through: ubiquity, perseverance, comprehensive annotation, automated training and certification, appropriate incentives, task engineering and variants of crowdsourcing. Building upon Linguistic Data Consortium's WebAnn framework, virtual front end web servers provide multiple interfaces to incentivize and engineer linguistic data contributions from targeted groups: linguists, citizen scientists, game players and students. Collection and annotation activities are analyzed into component tasks according to the skills they require and are assigned as appropriate to different workforces using different workflows. The combination of customized interfaces and novel incentive strategies enables ongoing, scalable data collection and annotation resulting in diverse language resources available to the wider Computer and Information Science and Engineering research and education communities.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Workshop on Sociolinguistic Archive Preparation
-
批准号:1144480
-
项目类别:Standard Grant
-
资助金额:$2.45万
-
财政年份:2011
-
负责人:Christopher Cieri
-
依托单位:
CRI:CRD Collaborative Research: General Techniques for Creating Treebanks with Multiple Representations: A Large-Scale Russian Application
-
批准号:0708276
-
项目类别:Standard Grant
-
资助金额:$1.25万
-
财政年份:2007
-
负责人:Christopher Cieri
-
依托单位:
Networking Data Centers
-
批准号:9982201
-
项目类别:Standard Grant
-
资助金额:$15.45万
-
财政年份:2000
-
负责人:Christopher Cieri
-
依托单位:
海外基金