EAGER: Mining a Year of Speech
EAGER: Mining a Year of Speech
批准号:
1048900
负责人:
Mark Liberman
金额:
$9.99万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2010
资助国家:
美国
项目状态:
已结题
起止时间:
2010-08-15 至 2012-07-31
中文摘要
存储和处理大量文本的技术是成熟且定义良好的。相比之下,从大量非文本材料(尤其是音频和视频)中浏览或挖掘内容的技术则不太发达。基于文本的大销售数据挖掘促进了相关学科的转型;处理口语的学科将从可访问的、可搜索的大型语料库中获得类似的好处。本项目探讨了为大量的美式英语和英式英语口语音频数据提供丰富、智能的数据挖掘功能的难题。它应用并扩展了最先进的技术,提供复杂、快速和灵活的访问,可以访问来自语言数据联盟、英国国家语料库和其他现有资源的丰富注释的一年语音语料库(约9,000小时,1亿单词或2 tb)。这比语音学、语言学和心理学等领域的研究人员以前使用的数据多10倍,是日常实践中使用的数据的100到1000倍。语音到文本的对齐和搜索工具将为许多领域的研究人员打开一个新的数据世界,从语言学和语音学到人类学,语音传播,口述历史和媒体研究。互联网上的音频视频使用量很大,并且以惊人的速度增长,提供了越来越多的材料,数量越来越多,范围越来越广。可靠的自动注释,索引和搜索这些材料将允许研究人员检查形式和内容的分布跨越时间,空间和社会结构。
英文摘要
Technologies for storing and processing vast amounts of text are mature and well-defined. In contrast, technologies for browsing or mining content from large collections of non-textual material, especially audio and video, are less well developed. Large sale data mining on text has helped transform the relevant disciplines; the disciplines dealing with spoken language will reap similar benefits from accessible, searchable, large corpora.This project explores the difficult problem of providing rich, intelligent data mining capabilities for a substantial collection of spoken audio data in American and British English. It applies and extends state-of-the-art techniques to offer sophisticated, rapid and flexible access to a richly annotated corpus of a year of speech (about 9,000 hours, 100 million words, or 2 terabytes), derived from the Linguistic Data Consortium, the British National Corpus, and other existing resources. This is ten times more data than has previously been used by researchers in fields such as phonetics, linguistics, and psychology, and 100 to 1,000 times the amounts that are used in common practice.Speech-to-text alignment and search tools will open a new universe of data to researchers in many fields, from linguistics and phonetics to anthropology, speech communication, oral history, and media studies. Audio-video usage on the internet is large and growing at an extraordinary rate, offering increasingly large amounts of an increasingly large range of material. Reliable automatic annotation, indexing and search of this material will allow researchers to examine the distribution of both form and content across time, space, and social structure.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CI-NEW: NIEUW: Novel Incentives and Workflows in Linguistic Data Collection and Annotation
-
批准号:1730377
-
项目类别:Standard Grant
-
资助金额:$121.85万
-
财政年份:2017
-
负责人:Mark Liberman
-
依托单位:
Language Preservation 2.0: Crowdsourcing Oral Language Documentation using Mobile Devices
-
批准号:1160639
-
项目类别:Standard Grant
-
资助金额:$10.15万
-
财政年份:2012
-
负责人:Mark Liberman
-
依托单位:
Prosodic Systems in New Guinea: Integrating computational and typological approaches to linguistic analysis
-
批准号:0951651
-
项目类别:Standard Grant
-
资助金额:$29.93万
-
财政年份:2010
-
负责人:Mark Liberman
-
依托单位:
Collaborative Research: OLAC: Accessing the World's Language Resources
-
批准号:0723357
-
项目类别:Continuing Grant
-
资助金额:$14.7万
-
财政年份:2007
-
负责人:Mark Liberman
-
依托单位:
ITR-SCOTUS: A Resource for Collaborative Research in Speech Technology, Linguistics, Decision Processes and the Law
-
批准号:0325739
-
项目类别:Continuing Grant
-
资助金额:$72.5万
-
财政年份:2003
-
负责人:Mark Liberman
-
依托单位:
Querying Linguistic Databases
-
批准号:0317826
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2003
-
负责人:Mark Liberman
-
依托单位:
Eletronic Materials For Natural Language Research
-
批准号:9113530
-
项目类别:Standard Grant
-
资助金额:$13.99万
-
财政年份:1991
-
负责人:Mark Liberman
-
依托单位:
国内基金
海外基金
基于Genome mining技术研究抑制表皮葡萄球菌生物膜形成的次级代谢产物
-
批准号:21242003
-
项目类别:专项基金项目
-
资助金额:10.0万元
-
批准年份:2012
-
负责人:昌军
-
依托单位: