课题基金 / 基金详情

CI-NEW: NIEUW: Novel Incentives and Workflows in Linguistic Data Collection and Annotation

CI-NEW: NIEUW: Novel Incentives and Workflows in Linguistic Data Collection and Annotation
CI-NEW:NIEUW:语言数据收集和注释中的新颖激励措施和工作流程
批准号:
1730377
负责人:
Mark Liberman
金额:
$121.85万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-07-15 至 2023-06-30

项目摘要

项目成果

Mark Liberman的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
Language touches every aspect of human life. People speak and write in order to manage relationships from the personal to the international, to gather and provide information, to negotiate, influence and inspire. Scientists use language to communicate their findings regardless of their field of study. Although researchers have been working for six decades to process language via computer, only in the past several years have their efforts have produced technologies of sufficient maturity that they can affect the lives of the average citizen. Today, some of the most fortunate use computers to search the vast archives of the Internet, to translate material from languages they do not understand into languages they do and to interact with smart devices by giving them natural language commands and queries and receive responses in kind. Despite the growth and promise of human language technologies, they are in fact available for only a tiny portion of the world's approximately 7000 languages and, even then, for only a limited range of situations. This is the case because the approaches that have proven most successful in developing human language technologies require vast amounts of spoken or written language material that have been augmented by human judgment as to their interpretation, but such resources are lacking for most languages and for many types of situations, even for languages of international importance, including English. This Research Infrastructure project will address this shortage of language resources by supporting the language technology research community to employ novel incentives and alternate workflows to greatly expand the methods that have been used to date for collecting and annotating language data. The resulting resources will support research and development on an expanded range of language technologies, leading to the creation and deployment of applications for an increasingly broad range of languages and situations. Even a brief observation of user behavior on social media, online games, citizen science and public good initiatives demonstrates that many people around the world are willing to devote collectively vast amounts of effort when given appropriate motivation and effective tools. This project will harness some of the immense people-power that drives such activities and focus it on problems of developing language resources that help computers learn to process language. Specifically, the project will create a software toolkit to be developed by the project team in response to the needs of language technology researchers to create online activities that yield language resources. The activities will include games, citizen science and tools for language professionals, clustered into a series of portals that appeal to different populations of users. The project will build and maintain the database and web servers, with redundancy, load balancing and fail over, to run the principal instance of all of the activities, and an open-source release of the software will enable other researchers to build their own instances independently. Finally, the data resulting from this project will be shared with the least restrictive terms possible to further support language technology research and development activities worldwide.
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
Using Games to Augment Corpora for Language Recognition and Confusability
使用游戏增强语料库以实现语言识别和混淆
DOI: 10.21437/interspeech.2021-1611
发表时间: 2021
期刊: Proceedings of the 22nd Annual Conference of the International Speech Communication Association
影响因子: --
作者: [Cieri, Christopher, Fiumara, James, Wright, Jonathan]
通讯作者: Wright, Jonathan
LanguageARC: Developing Language Resources Through Citizen Linguistics
LanguageARC:通过公民语言学开发语言资源
DOI: --
发表时间: 2020
期刊: Proceedings of the LREC 2020 Workshop Citizen Linguistics in Language Resource Development (CLLRD 2020
影响因子: --
作者: [Fiumara, James, Cieri, Christopher, Wright, Jonathan, Liberman, Mark]
通讯作者: Liberman, Mark
LanguageARC – a tutorial
LanguageARC — 教程
DOI: --
发表时间: 2020
期刊: Proceedings of the LREC 2020 Workshop Citizen Linguistics in Language Resource Development (CLLRD 2020
影响因子: --
作者: [Cieri, Christopher, Fiumara, James]
通讯作者: Fiumara, James
DOI: --
发表时间: 2022
期刊: Proceedings of the 13th Edition of the Language Resources and Evaluation Conference
影响因子: --
作者: [Christopher Cieri, Mark Liberman]
通讯作者: Christopher Cieri, Mark Liberman
Language Preservation 2.0: Crowdsourcing Oral Language Documentation using Mobile Devices
  • 批准号:
    1160639
  • 项目类别:
    Standard Grant
  • 资助金额:
    $10.15万
  • 财政年份:
    2012
  • 负责人:
    Mark Liberman
  • 依托单位:
EAGER: Mining a Year of Speech
  • 批准号:
    1048900
  • 项目类别:
    Standard Grant
  • 资助金额:
    $9.99万
  • 财政年份:
    2010
  • 负责人:
    Mark Liberman
  • 依托单位:
Prosodic Systems in New Guinea: Integrating computational and typological approaches to linguistic analysis
  • 批准号:
    0951651
  • 项目类别:
    Standard Grant
  • 资助金额:
    $29.93万
  • 财政年份:
    2010
  • 负责人:
    Mark Liberman
  • 依托单位:
Collaborative Research: OLAC: Accessing the World's Language Resources
  • 批准号:
    0723357
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $14.7万
  • 财政年份:
    2007
  • 负责人:
    Mark Liberman
  • 依托单位:
海外基金