课题基金 / 基金详情

CI-NEW: NIEUW: Novel Incentives and Workflows in Linguistic Data Collection and Annotation

CI-NEW: NIEUW: Novel Incentives and Workflows in Linguistic Data Collection and Annotation
CI-NEW:NIEUW:语言数据收集和注释中的新颖激励措施和工作流程
批准号:
1730377
负责人:
Mark Liberman
金额:
$121.85万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-07-15 至 2023-06-30

项目摘要

项目成果

Mark Liberman的其他基金

相似基金

相关文献

中文摘要
翻译
语言触及人类生活的方方面面。人们说话和写作是为了管理从个人到国际的关系,收集和提供信息,谈判,影响和激励。科学家们用语言来传达他们的发现,而不管他们的研究领域是什么。尽管研究人员用计算机处理语言已经工作了60年,但直到最近几年,他们的努力才产生了足够成熟的技术,可以影响普通公民的生活。今天,一些最幸运的人用电脑搜索互联网上的大量档案,把他们不懂的语言翻译成他们懂的语言,通过向智能设备发出自然语言命令和查询,与它们进行交互,并接收相应的回应。尽管人类语言技术不断发展,前景广阔,但事实上,它们只适用于世界上大约7000种语言中的一小部分,而且即便如此,也只适用于有限的情况。之所以会出现这种情况,是因为在开发人类语言技术方面被证明最成功的方法需要大量的口语或书面语材料,这些材料通过人类对其解释的判断而得到增强,但对于大多数语言和许多类型的情况,甚至对于包括英语在内的具有国际重要性的语言,都缺乏这种资源。该研究基础设施项目将通过支持语言技术研究社区采用新颖的激励措施和替代工作流程来极大地扩展迄今为止用于收集和注释语言数据的方法,从而解决语言资源的短缺问题。由此产生的资源将支持更多语言技术的研究和开发,从而为越来越广泛的语言和情况创建和部署应用程序。即使是对社交媒体、网络游戏、公民科学和公益倡议上的用户行为的简短观察也表明,如果有适当的动机和有效的工具,世界各地的许多人都愿意集体投入大量的努力。该项目将利用推动此类活动的巨大人力资源,并将其集中在开发语言资源以帮助计算机学习处理语言的问题上。具体来说,该项目将创建一个软件工具包,由项目团队开发,以响应语言技术研究人员创建产生语言资源的在线活动的需求。这些活动将包括游戏、公民科学和语言专业人士的工具,集中在一系列吸引不同用户群体的门户网站中。该项目将构建和维护数据库和web服务器,具有冗余,负载平衡和故障转移,以运行所有活动的主要实例,并且软件的开源版本将使其他研究人员能够独立构建自己的实例。最后,这个项目产生的数据将以尽可能少的限制条款共享,以进一步支持世界范围内的语言技术研究和开发活动。
英文摘要
Language touches every aspect of human life. People speak and write in order to manage relationships from the personal to the international, to gather and provide information, to negotiate, influence and inspire. Scientists use language to communicate their findings regardless of their field of study. Although researchers have been working for six decades to process language via computer, only in the past several years have their efforts have produced technologies of sufficient maturity that they can affect the lives of the average citizen. Today, some of the most fortunate use computers to search the vast archives of the Internet, to translate material from languages they do not understand into languages they do and to interact with smart devices by giving them natural language commands and queries and receive responses in kind. Despite the growth and promise of human language technologies, they are in fact available for only a tiny portion of the world's approximately 7000 languages and, even then, for only a limited range of situations. This is the case because the approaches that have proven most successful in developing human language technologies require vast amounts of spoken or written language material that have been augmented by human judgment as to their interpretation, but such resources are lacking for most languages and for many types of situations, even for languages of international importance, including English. This Research Infrastructure project will address this shortage of language resources by supporting the language technology research community to employ novel incentives and alternate workflows to greatly expand the methods that have been used to date for collecting and annotating language data. The resulting resources will support research and development on an expanded range of language technologies, leading to the creation and deployment of applications for an increasingly broad range of languages and situations. Even a brief observation of user behavior on social media, online games, citizen science and public good initiatives demonstrates that many people around the world are willing to devote collectively vast amounts of effort when given appropriate motivation and effective tools. This project will harness some of the immense people-power that drives such activities and focus it on problems of developing language resources that help computers learn to process language. Specifically, the project will create a software toolkit to be developed by the project team in response to the needs of language technology researchers to create online activities that yield language resources. The activities will include games, citizen science and tools for language professionals, clustered into a series of portals that appeal to different populations of users. The project will build and maintain the database and web servers, with redundancy, load balancing and fail over, to run the principal instance of all of the activities, and an open-source release of the software will enable other researchers to build their own instances independently. Finally, the data resulting from this project will be shared with the least restrictive terms possible to further support language technology research and development activities worldwide.
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
Using Games to Augment Corpora for Language Recognition and Confusability
使用游戏增强语料库以实现语言识别和混淆
DOI: 10.21437/interspeech.2021-1611
发表时间: 2021
期刊: Proceedings of the 22nd Annual Conference of the International Speech Communication Association
影响因子: --
作者: [Cieri, Christopher, Fiumara, James, Wright, Jonathan]
通讯作者: Wright, Jonathan
LanguageARC: Developing Language Resources Through Citizen Linguistics
LanguageARC:通过公民语言学开发语言资源
DOI: --
发表时间: 2020
期刊: Proceedings of the LREC 2020 Workshop Citizen Linguistics in Language Resource Development (CLLRD 2020
影响因子: --
作者: [Fiumara, James, Cieri, Christopher, Wright, Jonathan, Liberman, Mark]
通讯作者: Liberman, Mark
LanguageARC – a tutorial
LanguageARC — 教程
DOI: --
发表时间: 2020
期刊: Proceedings of the LREC 2020 Workshop Citizen Linguistics in Language Resource Development (CLLRD 2020
影响因子: --
作者: [Cieri, Christopher, Fiumara, James]
通讯作者: Fiumara, James
DOI: --
发表时间: 2022
期刊: Proceedings of the 13th Edition of the Language Resources and Evaluation Conference
影响因子: --
作者: [Christopher Cieri, Mark Liberman]
通讯作者: Christopher Cieri, Mark Liberman
Language Preservation 2.0: Crowdsourcing Oral Language Documentation using Mobile Devices
  • 批准号:
    1160639
  • 项目类别:
    Standard Grant
  • 资助金额:
    $10.15万
  • 财政年份:
    2012
  • 负责人:
    Mark Liberman
  • 依托单位:
EAGER: Mining a Year of Speech
  • 批准号:
    1048900
  • 项目类别:
    Standard Grant
  • 资助金额:
    $9.99万
  • 财政年份:
    2010
  • 负责人:
    Mark Liberman
  • 依托单位:
Prosodic Systems in New Guinea: Integrating computational and typological approaches to linguistic analysis
  • 批准号:
    0951651
  • 项目类别:
    Standard Grant
  • 资助金额:
    $29.93万
  • 财政年份:
    2010
  • 负责人:
    Mark Liberman
  • 依托单位:
Collaborative Research: OLAC: Accessing the World's Language Resources
  • 批准号:
    0723357
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $14.7万
  • 财政年份:
    2007
  • 负责人:
    Mark Liberman
  • 依托单位:
海外基金