课题基金 / 基金详情

SPeech Across Dialects of English (SPADE): large-scale digital analysis of a spoken language across space and time

SPeech Across Dialects of English (SPADE): large-scale digital analysis of a spoken language across space and time
Speech Across English Dialects of English (SPADE):跨空间和时间的口语大规模数字分析
批准号:
ES/R003963/1
负责人:
Jane StuartSmith
金额:
$20.51万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2017
资助国家:
英国
项目状态:
已结题
起止时间:
2017 至 --

项目摘要

项目成果

Jane StuartSmith的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
Obtaining a data visualization of a text search within seconds via generic, large-scale search algorithms, such as Google n-gram viewer, is available to anyone. By contrast, speech research is only now entering its own 'big data' revolution. Historically, linguistic research has tended to carry out fine-grained analysis of a few aspects of speech from one or a few languages or dialects. The current scale of speech research studies has shaped our understanding of spoken language and the kinds of questions that we ask. Today, massive digital collections of transcribed speech are available from many different languages, gathered for many different purposes: from oral histories, to large datasets for training speech recognition systems, to legal and political interactions. Sophisticated speech processing tools exist to analyze these data, but require substantial technical skill. Given this confluence of data and tools, linguists have a new opportunity to answer fundamental questions about the nature and development of spoken language. Our project seeks to establish the key tools to enable large-scale speech research to become as powerful and pervasive as large-scale text mining. It is based on a partnership of three teams based in Scotland, Canada and the US. Together we will exploit methods from computing science and put them to work with tools and methods from speech science, linguistics and digital humanities, to discover how much the sounds of English across the Atlantic vary over space and time. We will develop an innovative and user-friendly software which exploits the availability of existing speech data and speech processing tools to facilitate large-scale integrated speech corpus analysis across many datasets together. The gains of such an approach are substantial: linguists will be able to scale up answers to existing research questions from one to many varieties of a language, and ask new and different questions about spoken language within and across social, regional, and cultural, contexts. Computational linguistics, speech technology, forensic and clinical linguistics researchers, who engage with variability in spoken language, will also benefit directly from our software. This project will also open up vast potential for those who already use digital scholarship for spoken language collections in the humanities and social sciences more broadly, e.g. literary scholars, sociologists, anthropologists, historians, political scientists. The possibility of ethically non-invasive inspection of speech and texts will allow analysts to uncover far more than is possible through textual analysis alone.Our project will develop and apply our new software to a global language, English, using 43 existing public and private spoken datasets of Old World (British Isles) and New World (North American) English, across an effective time span of more than 100 years, spanning the entire 20th century. Much of what we know about spoken English comes from influential studies on a few specific aspects of speech from one or two dialects. This vast literature has established important research questions which can be investigated for the first time on a much larger scale, through standardized data across many different varieties of English. Our large-scale study will complement current-scale studies, by enabling us to consider stability and change in English across the 20th century on an unparalleled scale. The global nature of English means that our findings will be interesting and relevant to a large international non-academic audience; they will be made accessible through an innovative and dynamic visualization of linguistic variation via an interactive sound mapping website. In addition to new insights into spoken English, this project will also lay the crucial groundwork for large-scale speech studies across many datasets from different languages, of different formats and structures.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
The Open Handbook of Linguistic Data Management
语言数据管理开放手册
DOI: 10.7551/mitpress/12200.003.0020
发表时间: 2022
期刊:
影响因子: --
作者: [Sonderegger M]
通讯作者: Sonderegger M
DOI: 10.18653/v1/2022.sigmorphon-1.8
发表时间: 2022
期刊:
影响因子: --
作者: [Tanner J]
通讯作者: Tanner J
Age vectors vs. axes of intraspeaker variation in vowel formants measured automatically from several English speech corpora
从几个英语语音语料库自动测量的年龄向量与元音共振峰中说话者内部变化的轴
DOI: --
发表时间: 2019
期刊: International Congress of Phonetic Sciences
影响因子: --
作者: [Mielke, Jeff, Thomas, Erik R., Fruehwald, Josef, McAuliffe, Michael, Sonderegger, Morgan, Stuart-Smith, Jane, Dodsworth, Robin]
通讯作者: Dodsworth, Robin
Toward "English" Phonetics: Variability in the Pre-consonantal Voicing Effect Across English Dialects and Speakers.
朝向“英语”语音:英语方言和扬声器之间的辅助声音效果的变化。
DOI: 10.3389/frai.2020.00038
发表时间: 2020
期刊: Frontiers in artificial intelligence
影响因子: 4
作者: [Tanner J, Sonderegger M, Stuart-Smith J, Fruehwald J]
通讯作者: Fruehwald J
9
    Dynamic dialects: integrating articulatory video to reveal the complexity of speech
    • 批准号:
      AH/L010380/1
    • 项目类别:
      Research Grant
    • 资助金额:
      $23.69万
    • 财政年份:
      2014
    • 负责人:
      Jane StuartSmith
    • 依托单位:
    国内基金
    海外基金
    基于鱼血模型研究几种典型人用药物的Read-across假设
    • 批准号:
      21577103
    • 项目类别:
      面上项目
    • 资助金额:
      65.0万元
    • 批准年份:
      2015
    • 负责人:
      胡霞林
    • 依托单位: