EAGER: Mining a Year of Speech
EAGER: Mining a Year of Speech
批准号:
1048900
负责人:
Mark Liberman
金额:
$9.99万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2010
资助国家:
美国
项目状态:
已结题
起止时间:
2010-08-15 至 2012-07-31
中文摘要
存储和处理大量文本的技术已经成熟且定义明确。相比之下,用于从大量非文本材料(特别是音频和视频)中浏览或挖掘内容的技术则不太发达。文本的大规模销售数据挖掘已经帮助相关学科发生了转变;处理口语的学科也将从可访问、可搜索的大型语料库中获得类似的好处。本项目探讨了为大量的美国和英国英语口语音频数据提供丰富、智能的数据挖掘能力的难题。 它应用并扩展了最先进的技术,以提供对一年语音(约9,000小时,1亿字或2 TB)的丰富注释语料库的复杂,快速和灵活的访问,这些语料库来自语言数据联盟,英国国家语料库和其他现有资源。这比语音学、语言学和心理学等领域的研究人员以前使用的数据多10倍,是通常使用的数据量的100到1,000倍。语音到文本对齐和搜索工具将为许多领域的研究人员打开一个新的数据世界,从语言学和语音学到人类学、语音传播、口述历史和媒体研究。互联网上的音频视频使用量很大,并且以惊人的速度增长,提供越来越多的材料。可靠的自动注释,索引和搜索这些材料将使研究人员能够检查跨时间,空间和社会结构的形式和内容的分布。
英文摘要
Technologies for storing and processing vast amounts of text are mature and well-defined. In contrast, technologies for browsing or mining content from large collections of non-textual material, especially audio and video, are less well developed. Large sale data mining on text has helped transform the relevant disciplines; the disciplines dealing with spoken language will reap similar benefits from accessible, searchable, large corpora.This project explores the difficult problem of providing rich, intelligent data mining capabilities for a substantial collection of spoken audio data in American and British English. It applies and extends state-of-the-art techniques to offer sophisticated, rapid and flexible access to a richly annotated corpus of a year of speech (about 9,000 hours, 100 million words, or 2 terabytes), derived from the Linguistic Data Consortium, the British National Corpus, and other existing resources. This is ten times more data than has previously been used by researchers in fields such as phonetics, linguistics, and psychology, and 100 to 1,000 times the amounts that are used in common practice.Speech-to-text alignment and search tools will open a new universe of data to researchers in many fields, from linguistics and phonetics to anthropology, speech communication, oral history, and media studies. Audio-video usage on the internet is large and growing at an extraordinary rate, offering increasingly large amounts of an increasingly large range of material. Reliable automatic annotation, indexing and search of this material will allow researchers to examine the distribution of both form and content across time, space, and social structure.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CI-NEW: NIEUW: Novel Incentives and Workflows in Linguistic Data Collection and Annotation
-
批准号:1730377
-
项目类别:Standard Grant
-
资助金额:$121.85万
-
财政年份:2017
-
负责人:Mark Liberman
-
依托单位:
Language Preservation 2.0: Crowdsourcing Oral Language Documentation using Mobile Devices
-
批准号:1160639
-
项目类别:Standard Grant
-
资助金额:$10.15万
-
财政年份:2012
-
负责人:Mark Liberman
-
依托单位:
Prosodic Systems in New Guinea: Integrating computational and typological approaches to linguistic analysis
-
批准号:0951651
-
项目类别:Standard Grant
-
资助金额:$29.93万
-
财政年份:2010
-
负责人:Mark Liberman
-
依托单位:
Collaborative Research: OLAC: Accessing the World's Language Resources
-
批准号:0723357
-
项目类别:Continuing Grant
-
资助金额:$14.7万
-
财政年份:2007
-
负责人:Mark Liberman
-
依托单位:
ITR-SCOTUS: A Resource for Collaborative Research in Speech Technology, Linguistics, Decision Processes and the Law
-
批准号:0325739
-
项目类别:Continuing Grant
-
资助金额:$72.5万
-
财政年份:2003
-
负责人:Mark Liberman
-
依托单位:
Querying Linguistic Databases
-
批准号:0317826
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2003
-
负责人:Mark Liberman
-
依托单位:
Eletronic Materials For Natural Language Research
-
批准号:9113530
-
项目类别:Standard Grant
-
资助金额:$13.99万
-
财政年份:1991
-
负责人:Mark Liberman
-
依托单位:
国内基金
海外基金
基于Genome mining技术研究抑制表皮葡萄球菌生物膜形成的次级代谢产物
-
批准号:21242003
-
项目类别:专项基金项目
-
资助金额:10.0万元
-
批准年份:2012
-
负责人:昌军
-
依托单位: