课题基金 / 基金详情

SPeech Across Dialects of English (SPADE): large-scale digital analysis of a spoken language across space and time

SPeech Across Dialects of English (SPADE): large-scale digital analysis of a spoken language across space and time
Speech Across English Dialects of English (SPADE):跨空间和时间的口语大规模数字分析
批准号:
ES/R003963/1
负责人:
Jane StuartSmith
金额:
$20.51万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2017
资助国家:
英国
项目状态:
已结题
起止时间:
2017 至 --

项目摘要

项目成果

Jane StuartSmith的其他基金

相似基金

相关文献

中文摘要
翻译
任何人都可以通过通用的大规模搜索算法(如Google n-gram查看器)在几秒钟内获得文本搜索的数据可视化。相比之下,语音研究现在才进入自己的“大数据”革命。从历史上看,语言学研究倾向于对一种或几种语言或方言的几个方面进行细粒度的分析。目前的言语研究规模塑造了我们对口语的理解和我们提出的问题的种类。今天,从口述历史到用于培训语音识别系统的大型数据集,再到法律和政治互动,人们可以从许多不同的语言获得大量转录语音的数字集合,用于许多不同的目的。有复杂的语音处理工具可以分析这些数据,但需要大量的技术技能。鉴于数据和工具的这种融合,语言学家有了一个新的机会来回答关于口语的性质和发展的基本问题。我们的项目旨在建立关键工具,使大规模语音研究变得像大规模文本挖掘一样强大和普遍。它基于苏格兰、加拿大和美国的三支球队的合作伙伴关系。我们将共同利用计算科学的方法,并将它们与言语科学、语言学和数字人文科学的工具和方法结合起来,以发现大西洋彼岸英语的发音在空间和时间上的差异。我们将开发一个创新和用户友好的软件,利用现有的语音数据和语音处理工具,促进对许多数据集的大规模综合语音语料库分析。这种方法的收获是巨大的:语言学家将能够从一种语言的一种变体扩大到多种语言变体的现有研究问题的答案,并在社会、地区和文化背景下提出关于口语的新的不同问题。从事口语可变性研究的计算语言学、语音技术、法医和临床语言学研究人员也将直接受益于我们的软件。该项目还将为那些已经在更广泛的人文和社会科学领域利用数字奖学金收集口语的人,如文学学者、社会学家、人类学家、历史学家和政治学家,带来巨大的潜力。对语音和文本进行伦理非侵入性检查的可能性将使分析师能够发现远远超过仅通过文本分析所能发现的内容。我们的项目将使用43个现有的旧世界(不列颠群岛)和新世界(北美)英语的公共和私人口语数据库,开发并将我们的新软件应用于全球语言英语,有效时间跨度超过100年,跨越整个20世纪。我们对英语口语的许多了解都来自于对一两种方言的几个特定方面进行的有影响力的研究。这些浩瀚的文献已经确立了重要的研究问题,这些问题可以通过许多不同英语变体的标准化数据首次在更大范围内进行调查。我们的大规模研究将补充目前的研究,使我们能够以前所未有的规模考虑英语在20世纪的稳定性和变化。英语的全球性意味着我们的调查结果将是有趣的,并与大量的国际非学术受众相关;这些调查结果将通过交互式声音地图网站以创新和动态的语言变化可视化方式提供。除了对英语口语的新见解外,该项目还将为大规模语音研究奠定关键基础,这些数据来自不同语言、不同格式和结构的许多数据集。
英文摘要
Obtaining a data visualization of a text search within seconds via generic, large-scale search algorithms, such as Google n-gram viewer, is available to anyone. By contrast, speech research is only now entering its own 'big data' revolution. Historically, linguistic research has tended to carry out fine-grained analysis of a few aspects of speech from one or a few languages or dialects. The current scale of speech research studies has shaped our understanding of spoken language and the kinds of questions that we ask. Today, massive digital collections of transcribed speech are available from many different languages, gathered for many different purposes: from oral histories, to large datasets for training speech recognition systems, to legal and political interactions. Sophisticated speech processing tools exist to analyze these data, but require substantial technical skill. Given this confluence of data and tools, linguists have a new opportunity to answer fundamental questions about the nature and development of spoken language. Our project seeks to establish the key tools to enable large-scale speech research to become as powerful and pervasive as large-scale text mining. It is based on a partnership of three teams based in Scotland, Canada and the US. Together we will exploit methods from computing science and put them to work with tools and methods from speech science, linguistics and digital humanities, to discover how much the sounds of English across the Atlantic vary over space and time. We will develop an innovative and user-friendly software which exploits the availability of existing speech data and speech processing tools to facilitate large-scale integrated speech corpus analysis across many datasets together. The gains of such an approach are substantial: linguists will be able to scale up answers to existing research questions from one to many varieties of a language, and ask new and different questions about spoken language within and across social, regional, and cultural, contexts. Computational linguistics, speech technology, forensic and clinical linguistics researchers, who engage with variability in spoken language, will also benefit directly from our software. This project will also open up vast potential for those who already use digital scholarship for spoken language collections in the humanities and social sciences more broadly, e.g. literary scholars, sociologists, anthropologists, historians, political scientists. The possibility of ethically non-invasive inspection of speech and texts will allow analysts to uncover far more than is possible through textual analysis alone.Our project will develop and apply our new software to a global language, English, using 43 existing public and private spoken datasets of Old World (British Isles) and New World (North American) English, across an effective time span of more than 100 years, spanning the entire 20th century. Much of what we know about spoken English comes from influential studies on a few specific aspects of speech from one or two dialects. This vast literature has established important research questions which can be investigated for the first time on a much larger scale, through standardized data across many different varieties of English. Our large-scale study will complement current-scale studies, by enabling us to consider stability and change in English across the 20th century on an unparalleled scale. The global nature of English means that our findings will be interesting and relevant to a large international non-academic audience; they will be made accessible through an innovative and dynamic visualization of linguistic variation via an interactive sound mapping website. In addition to new insights into spoken English, this project will also lay the crucial groundwork for large-scale speech studies across many datasets from different languages, of different formats and structures.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
The Open Handbook of Linguistic Data Management
语言数据管理开放手册
DOI: 10.7551/mitpress/12200.003.0020
发表时间: 2022
期刊:
影响因子: --
作者: [Sonderegger M]
通讯作者: Sonderegger M
DOI: 10.18653/v1/2022.sigmorphon-1.8
发表时间: 2022
期刊:
影响因子: --
作者: [Tanner J]
通讯作者: Tanner J
Age vectors vs. axes of intraspeaker variation in vowel formants measured automatically from several English speech corpora
从几个英语语音语料库自动测量的年龄向量与元音共振峰中说话者内部变化的轴
DOI: --
发表时间: 2019
期刊: International Congress of Phonetic Sciences
影响因子: --
作者: [Mielke, Jeff, Thomas, Erik R., Fruehwald, Josef, McAuliffe, Michael, Sonderegger, Morgan, Stuart-Smith, Jane, Dodsworth, Robin]
通讯作者: Dodsworth, Robin
Toward "English" Phonetics: Variability in the Pre-consonantal Voicing Effect Across English Dialects and Speakers.
朝向“英语”语音:英语方言和扬声器之间的辅助声音效果的变化。
DOI: 10.3389/frai.2020.00038
发表时间: 2020
期刊: Frontiers in artificial intelligence
影响因子: 4
作者: [Tanner J, Sonderegger M, Stuart-Smith J, Fruehwald J]
通讯作者: Fruehwald J
共 9 条
    Dynamic dialects: integrating articulatory video to reveal the complexity of speech
    • 批准号:
      AH/L010380/1
    • 项目类别:
      Research Grant
    • 资助金额:
      $23.69万
    • 财政年份:
      2014
    • 负责人:
      Jane StuartSmith
    • 依托单位:
    国内基金
    海外基金
    基于鱼血模型研究几种典型人用药物的Read-across假设
    • 批准号:
      21577103
    • 项目类别:
      面上项目
    • 资助金额:
      65.0万元
    • 批准年份:
      2015
    • 负责人:
      胡霞林
    • 依托单位: