课题基金 / 基金详情

EAGER: Mining a Year of Speech

EAGER: Mining a Year of Speech
EAGER:挖掘一年的演讲
批准号:
1048900
负责人:
Mark Liberman
金额:
$9.99万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2010
资助国家:
美国
项目状态:
已结题
起止时间:
2010-08-15 至 2012-07-31
关键词:

项目摘要

项目成果

Mark Liberman的其他基金

相似基金

相关文献

中文摘要
翻译
存储和处理大量文本的技术已经成熟且定义良好。相比之下,从大量非文本材料,特别是音频和视频中浏览或挖掘内容的技术还不太发达。基于文本的大规模销售数据挖掘帮助改变了相关学科;涉及口语的学科将从可访问、可搜索的大型语料库中获得类似的好处。该项目探索了为大量美式和英式英语口语音频数据集合提供丰富、智能的数据挖掘功能这一难题。它应用并扩展了最先进的技术,以提供复杂、快速和灵活的访问一年语音(约9000小时、1亿字或2TB)的丰富注释语料库,这些语料库来自语言数据联盟、英国国家语料库和其他现有资源。这是语音学、语言学和心理学等领域研究人员以前使用的数据量的十倍,是一般实践中使用的数据量的100到1000倍。语音到文本对齐和搜索工具将为从语言学和语音学到人类学、言语交际、口述历史和媒体研究等许多领域的研究人员打开一个新的数据宇宙。互联网上的音视频使用量很大,并且以惊人的速度增长,提供了越来越多的材料,范围越来越广。对这些材料进行可靠的自动注释、索引和搜索将使研究人员能够检查形式和内容在时间、空间和社会结构上的分布。
英文摘要
Technologies for storing and processing vast amounts of text are mature and well-defined. In contrast, technologies for browsing or mining content from large collections of non-textual material, especially audio and video, are less well developed. Large sale data mining on text has helped transform the relevant disciplines; the disciplines dealing with spoken language will reap similar benefits from accessible, searchable, large corpora.This project explores the difficult problem of providing rich, intelligent data mining capabilities for a substantial collection of spoken audio data in American and British English. It applies and extends state-of-the-art techniques to offer sophisticated, rapid and flexible access to a richly annotated corpus of a year of speech (about 9,000 hours, 100 million words, or 2 terabytes), derived from the Linguistic Data Consortium, the British National Corpus, and other existing resources. This is ten times more data than has previously been used by researchers in fields such as phonetics, linguistics, and psychology, and 100 to 1,000 times the amounts that are used in common practice.Speech-to-text alignment and search tools will open a new universe of data to researchers in many fields, from linguistics and phonetics to anthropology, speech communication, oral history, and media studies. Audio-video usage on the internet is large and growing at an extraordinary rate, offering increasingly large amounts of an increasingly large range of material. Reliable automatic annotation, indexing and search of this material will allow researchers to examine the distribution of both form and content across time, space, and social structure.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CI-NEW: NIEUW: Novel Incentives and Workflows in Linguistic Data Collection and Annotation
  • 批准号:
    1730377
  • 项目类别:
    Standard Grant
  • 资助金额:
    $121.85万
  • 财政年份:
    2017
  • 负责人:
    Mark Liberman
  • 依托单位:
Language Preservation 2.0: Crowdsourcing Oral Language Documentation using Mobile Devices
  • 批准号:
    1160639
  • 项目类别:
    Standard Grant
  • 资助金额:
    $10.15万
  • 财政年份:
    2012
  • 负责人:
    Mark Liberman
  • 依托单位:
Prosodic Systems in New Guinea: Integrating computational and typological approaches to linguistic analysis
  • 批准号:
    0951651
  • 项目类别:
    Standard Grant
  • 资助金额:
    $29.93万
  • 财政年份:
    2010
  • 负责人:
    Mark Liberman
  • 依托单位:
Collaborative Research: OLAC: Accessing the World's Language Resources
  • 批准号:
    0723357
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $14.7万
  • 财政年份:
    2007
  • 负责人:
    Mark Liberman
  • 依托单位:
国内基金
海外基金
基于Genome mining技术研究抑制表皮葡萄球菌生物膜形成的次级代谢产物
  • 批准号:
    21242003
  • 项目类别:
    专项基金项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2012
  • 负责人:
    昌军
  • 依托单位: