课题基金 / 基金详情

Computational Methods for Speech Analysis

Computational Methods for Speech Analysis
语音分析的计算方法
批准号:
2120087
负责人:
Christopher Lucas
金额:
$24.93万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2021
资助国家:
美国
项目状态:
已结题
起止时间:
2021-08-01 至 2024-07-31

项目摘要

项目成果

Christopher Lucas的其他基金

相似基金

相关文献

中文摘要
翻译
该研究项目将开发用于测试人类交流假设的工具。研究人员通常从省略了声调的文本记录中研究人类交流。该项目将直接解决数据生成过程(其中扬声器和听众使用听觉通道来传达文本和非文本信号)与丢弃语音音频的普遍做法之间的脱节。研究人员将扩展他们以前的语音模型,音频和语音结构模型,以解决该模型的一些局限性。特别是,统计扩展将适应多个扬声器,并允许文本和音调的联合建模。为了证明统计扩展的价值,该模型将被应用到两个原始的视频语料库-警察的身体磨损的摄像机镜头和竞选演讲的联邦办公室。将开发新的软件,使研究人员能够轻松地快速注释大量的语音音频。基于浏览器的工具将支持自动和手动分割,沿着标签。多名研究生将获得计算密集型研究和软件开发方面的经验。该研究项目将扩展音频和语音结构模型(MASS),该模型将对话分析为嵌套的随机过程,其中(i)对话流展开为基于上下文协变量的说话者及其音调之间的一系列话语转换;以及(ii)每个话语内的听觉信号展开为在生成声音的音素之间转换的隐马尔可夫模型。该模型使社会科学家能够测试关于对话如何由固定协变量(例如,说话者性别,对话角色)和时变协变量(例如,外生的外部刺激、内生的对话轨迹(诸如先前说话者的语气)。然而,在其当前的实现中,MASS有两个关键的限制:首先,它使用资源密集型的人类注释每个扬声器的音调,这限制了应用程序与许多独特的扬声器,如警察身体佩戴的摄像机镜头的上下文。该项目将开发扩展,允许模型通过部分汇集具有相似语音特征的扬声器来借用力量。其次,MASS将文本作为外部给定的元数据。该项目将开发一种新的方法,用于文本和音频的联合建模,该方法将动态主题模型纳入MASS的会话流层。调查人员将进行两个应用程序,以证明多扬声器和联合文本音频建模扩展的价值。该奖项反映了NSF的法定使命,并已被认为是值得通过使用基金会的智力价值和更广泛的影响审查标准进行评估的支持。
英文摘要
This research project will develop tools for testing hypotheses about human communication. Researchers generally study human communication from textual transcripts which omit vocal tone. The project will directly address the disconnect between the data-generating process - in which speakers and listeners use the auditory channel to convey both textual and non-textual signals - and the widespread practice of discarding speech audio. The investigators will extend their prior speech model, The Model of Audio and Speech Structure, to address some limitations of the model. In particular, the statistical extensions will accommodate multiple speakers and allow for the joint modeling of text and tone. To demonstrate the value of the statistical extensions, the model will be applied to two original video corpora - police body-worn camera footage and campaign speeches for federal office. New software will be developed that makes it easy for researchers to quickly annotate a large amount of speech audio. The browser-based tools will enable automatic and manual segmentation, along with labeling. Multiple graduate students will gain experience in computationally intensive research and software development. The tools to be developed will be incorporated into ongoing public-private collaborations to improve oversight of police officers in the field.This research project will extend the Model of Audio and Speech Structure (MASS), which analyzes conversation as a nested stochastic process in which (i) the flow of conversation unfolds as a sequence of utterances transitioning between speakers and their vocal tones, based on contextual covariates; and (ii) the auditory signal within each utterance unfolds as a hidden Markov model that transitions between phonemes which generate sound. The model enables social scientists to test hypotheses about how conversations are structured by fixed covariates (e.g., speaker gender, conversation role) and time-varying covariates (e.g., exogenous external stimuli, endogenous conversation trajectory such as the previous speaker's tone). In its current implementation, however, MASS has two key limitations: First, it uses resource-intensive human annotations of tone for each speaker, which limits application to contexts with many unique speakers, such as police body-worn camera footage. This project will develop extensions allowing the model to borrow strength by partial pooling across speakers with similar speech profiles. Second, MASS incorporates text as externally given metadata. The project will develop a new approach for joint modeling of text and audio which will incorporate a dynamic topic model into the flow-of-conversation layer of MASS. The investigators will conduct two applications to demonstrate the value of the multi-speaker and joint text-audio modeling extensions.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
XMaS: The National Material Science Beamline Research Facility at the ESRF
  • 批准号:
    EP/Y031164/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $475.99万
  • 财政年份:
    2024
  • 负责人:
    Christopher Lucas
  • 依托单位:
Dissecting macrophage regulation of lung epithelial regeneration
  • 批准号:
    MR/X019314/1
  • 项目类别:
    Fellowship
  • 资助金额:
    $238.26万
  • 财政年份:
    2023
  • 负责人:
    Christopher Lucas
  • 依托单位:
XMaS Capital Equipment Upgrade
  • 批准号:
    EP/X035131/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $55.39万
  • 财政年份:
    2023
  • 负责人:
    Christopher Lucas
  • 依托单位:
Inflammation in Covid-19: Exploration of Critical Aspects of Pathogenesis (ICECAP)
  • 批准号:
    MR/V028790/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $50.58万
  • 财政年份:
    2020
  • 负责人:
    Christopher Lucas
  • 依托单位:
国内基金
海外基金
Computational Methods for Analyzing Toponome Data