课题基金 / 基金详情

STREAMLInED: Shared Tasks for Rapid, Efficient Analysis of Many Languages in Emerging Documentation

STREAMLInED: Shared Tasks for Rapid, Efficient Analysis of Many Languages in Emerging Documentation
STREAMLInED:用于快速、高效分析新兴文档中多种语言的共享任务
批准号:
1760475
负责人:
Gina-Anne Levow
金额:
$12.5万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-06-15 至 2024-09-30

项目摘要

项目成果

Gina-Anne Levow的其他基金

相似基金

相关文献

中文摘要
翻译
该项目将两个不同的科学和工程团体的研究兴趣结合在一起,以推动自动语音处理技术的边界,并将其好处带到濒危语言文档的紧急任务中。通过自动字幕和语音驱动的个人助理等工具,自动语音处理技术已经在许多说英语和其他广泛使用语言的人的日常生活中变得熟悉。与此同时,语言学家正争先恐后地记录和分析数以千计的语言,这些语言到本世纪末将不再被儿童习得。对记录的濒危语言口语数据进行自动处理将极大地帮助这项工作。然而,现代自动语音处理工具需要的训练数据集比濒危语言的可用数据集大几个数量级。该项目将通过围绕基于语言文件的数据集构建一个“共同任务评估挑战”来促进关于这一问题的科学知识。更好的语言文件使社区能够更好地进行语言振兴,这反过来又可以成为边缘化人口社区发展的一个关键组成部分。更广泛的影响还包括将处理小数据集的语音技术引入广泛使用但研究不足的语言,通常是对国家利益具有地缘政治和经济重要性的地区的交流语言。语言文档项目通常从大量录制的语音开始。将口语信号转换为转录形式是语言文件编制过程中的一个主要瓶颈。同样,语言档案馆保存着来自许多语言的记录的、未经分析的数据,这些语言没有活着的流利的语言,但有社区有兴趣重振其遗产语言。与此同时,开发能够在非常小的训练数据集上有效工作的技术对语音研究人员来说是一个开放而有趣的挑战。共同任务评估挑战框架提供了一种友好竞争的结构,不同的研究小组可以探索和比较使用标准化数据和指标进行评估的方法。几十年来,这种集中研究努力的战略推动了语言技术的前沿。该项目将首次将其应用于濒危语言记录的具体挑战:使用真正低资源的语言,通常有噪音或其他不完美的记录条件。挑战将集中于的具体任务包括:识别每段录音的语言和说话者,识别录音片段的体裁(例如,讲故事与对话),以及将简短的部分转录与口语录音对齐。研究人员将准备数据(根据语文档案中确定的现有数据集),建立任务参与者可用于比较和(或)进一步建立的有效基线系统,建立评价指标,并执行共同任务。共享任务结构将鼓励和支持参与者将其投稿开放源码,以确保向语文文件研究人员提供这些投稿。该项目还将包括与语言文献社区的接触,以培训这些研究人员使用所开发的技术。这一奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This project aligns the research interests of two separate scientific and engineering communities in order to push the boundaries of automatic speech processing technology and bring its benefits to the urgent task of endangered language documentation. Automatic speech processing technology has become familiar in the everyday lives of many speakers of English and other widely spoken languages through tools such as automatic captioning and voice-driven personal assistants. Meanwhile, linguists are rushing to document and analyze the thousands of languages that by the end of this century, will no longer be acquired by children. Such work would be greatly assisted by automatic processing of recorded spoken endangered language data. Modern automatic speech processing tools, however, require training data sets orders of magnitude larger than what is available for endangered languages. This project will advance scientific knowledge on this problem by structuring a "shared task evaluation challenge" around language documentation-based data sets. Better language documentation puts communities in a better position to undertake language revitalization, which in turn can be a key component of community development for marginalized populations. Broader impacts also include the benefits of bringing speech technology that works with small data sets to widely spoken but understudied languages, often languages of communication in regions of geopolitical and economic importance to national interests. Language documentation projects typically begin with large quantities of recorded speech. Turning that spoken signal into a transcribed form is a major bottleneck in the language documentation process. Similarly, language archives house recorded, unanalyzed data from many languages with no living fluent speaker, but which have communities interested in revitalizing their heritage languages. At the same time, the development of technology that can work effectively with very small training data sets is an open and interesting challenge for speech researchers. The shared task evaluation challenge framework provides the structure of a friendly competition in which different research groups can explore and compare approaches that are evaluated with standardized data and metrics. This strategy for focusing research effort has advanced the frontiers of language technology for decades. This project will apply it for the first time to the specific challenges of endangered language documentation: working with truly low-resource languages, with often noisy or other imperfect recording conditions. The specific tasks the challenge will focus on include: identifying the language and speaker of each segment of a recording, identifying the genre (e.g. story telling vs. dialogue) of segments of recordings, and aligning short partial transcriptions to the spoken recordings. The researchers will prepare the data (based on existing data sets identified in language archives), set up functioning baseline systems that task participants can use for comparison and/or build on further, establish evaluation metrics, and execute the shared task. The shared task structure will encourage and support participants in making their contributions open source, with an eye towards ensuring they are available to language documentation researchers. The project will also include outreach to the language documentation community in order to train such researchers in the use of the technology developed.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
DOI: --
发表时间: 2023
期刊:
影响因子: --
作者: [Gina-Anne Levow]
通讯作者: Gina-Anne Levow
DOI: 10.33011/computel.v1i.967
发表时间: 2021-03
期刊: Proceedings of the Workshop on Computational Methods for Endangered Languages
影响因子: --
作者: [Gina-Anne Levow;Emily Ahn;Emily M. Bender]
通讯作者: Gina-Anne Levow;Emily Ahn;Emily M. Bender
EL-STEC: Shared Task Evaluation Campaigns with Endangered Language Data
  • 批准号:
    1500157
  • 项目类别:
    Standard Grant
  • 资助金额:
    $5.0万
  • 财政年份:
    2015
  • 负责人:
    Gina-Anne Levow
  • 依托单位:
EAGER: ATAROS: Automatic Tagging and Recognition of Stance
  • 批准号:
    1351034
  • 项目类别:
    Standard Grant
  • 资助金额:
    $25.0万
  • 财政年份:
    2013
  • 负责人:
    Gina-Anne Levow
  • 依托单位:
Learning Tone
  • 批准号:
    0414919
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $0.0万
  • 财政年份:
    2004
  • 负责人:
    Gina-Anne Levow
  • 依托单位:
海外基金