RAPID: Harvesting Speech Datasets for Linguistic Research on the Web (Digging into Data Challenge)
RAPID: Harvesting Speech Datasets for Linguistic Research on the Web (Digging into Data Challenge)
批准号:
1035151
负责人:
Mats Rooth
金额:
$10.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2010
资助国家:
美国
项目状态:
已结题
起止时间:
2010-08-01 至 2013-07-31
中文摘要
韵律差异(节奏、重音和语调)在口语中普遍存在。对于以英语为母语的人来说,什么韵律在特定的句子和语境中最合适似乎是显而易见的,语言学和相关领域的研究人员对此提出了许多形式化的假设。但要确定这些假说的真实性是非常难以捉摸的。很大程度上的问题是,很难观察到足够多的给定现象的例子来评估假说。该项目旨在通过从网络资源中收集或“收获”特定单词序列或单词模式的实例来解决数据匮乏的问题。经常可以找到成百上千个使用完全相同的单词模式的人的例子。如果这些例子被收集到一个数据集,并提供给研究界,就有可能以前所未有的规模评估有关韵律形式和意义的理论。扩大现有数据的规模有望对我们对韵律的理解产生革命性的影响。口语的音频和音频-视频记录,包括播客、广播和电视广播、讲座和许多其他内容,在网络上无处不在。这本身没有帮助,因为不可能为了找到某一类型的几百个例子而收听数万个小时的演讲。幸运的是,越来越多的网站提供通过自动语音识别获得的文本转录(例如,福克斯商业新闻、WNYC、谷歌的选举视频搜索和麻省理工学院的大学讲座)。行业博客和时事通讯表明,很快会有更多大型网站上线。通过在文本转录中搜索单词模式并随后检索音频或视频文件,可以找到相关数据。为了从这些网络资源中构建韵律研究的数据集,项目团队将实现通过标准协议与网络交互的软件收获引擎。将收集8到12个特定现象的数据集。为了展示数据密集型方法的影响,将使用统计学和形式语言学的技术对样本进行分析。例如,一种称为机器学习分类的方法将被用来识别负责韵律感知的声音信号的特定特征(如音高、元音持续时间和强度)。韵律和语调在使语篇连贯、在传达信息的哪一部分是前景化和背景化以及消除说话者意图方面起着重要作用。韵律理解的任何进步不仅将加深我们对人类语言能力的理解,而且还将在语言教学、翻译研究、言语治疗、提高合成语音的可理解性以及改进语音识别系统等方面产生广泛的影响。
英文摘要
Distinctions of prosody (rhythm, stress, and intonation) are ubiquitous in spoken language. It often seems obvious to a native speakers of English what prosody is most appropriate in a given sentence and context, and researchers in Linguistics and related fields have proposed numerous formalized hypotheses about it. But establishing the validity of these hypotheses is remarkably elusive. Much of the problem is that it is difficult to observe enough examples of a given phenomenon to evaluate hypotheses. The project aims to address this problem of a dearth of data by collecting or "harvesting" examples of specific word sequences or word patterns from web sources. It is often possible to find hundreds or thousands of examples of people using the very same word pattern. If these examples are collected together into a dataset and made available to the research community, it will be possible to evaluate theories about the form and meaning of prosody on an unprecedented scale. Scaling up available data can be expected to have a transformative effect on our understanding of prosody.Audio and audio-video recordings of spoken language, including podcasts, radio and television broadcasts, lectures, and much else, are pervasive on the web. This does not help in itself, because it is not possible to listen to tens of thousands of hours of speech in order to find a few hundred examples of a certain type. Fortunately, more sites are becoming available that provide text transcriptions obtained with automatic speech recognition (for instance Fox Business News, WNYC, Elections Video Search at Google, and university lectures at MIT). Industry blogs and newsletters indicate that more large sites will come online soon. By searching for a word pattern in the text transcription and subsequently retrieving an audio or video file, it becomes possible to find relevant data. To construct datasets for prosody research from these web sources, the project team will implement software harvest engines that interact with the web through standard protocols. Datasets for eight to twelve specific phenomena will be collected. In order to demonstrate the impact of a data-intensive methodology, the samples will be analyzed using techniques of statistics and formal linguistics. For instance, an approach known as machine learning classification will be used to identify the specific features of the sound signal (such as pitch, vowel duration, and intensity) that are responsible for the perception of prosody.Prosody and intonation play an important role in making the discourse coherent, in signaling what part of the communicated information is foregrounded and backgrounded, and disambiguating speaker intention. Any advancement in understanding prosody will not only deepen our understanding of the human language capability, it also has implications in a wide range of areas, including language instruction, translation studies, speech therapy, improving comprehensibility of synthesized speech, and improving speech recognition systems.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金